
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Quick Answer
Google DeepMind has launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, enhancing AI agent efficiency with 17% lower token usage in 3.6 Flash and 350 tokens/sec in 3.5 Flash-Lite.
Quick Take
These models improve performance metrics across various benchmarks, making them more cost-effective for developers.
Key Points
- 3.6 Flash reduces output token usage by 17% compared to 3.5 Flash.
- 3.5 Flash-Lite achieves 350 output tokens per second, enhancing throughput.
- 3.6 Flash shows improved precision in coding tasks, outperforming benchmarks.
- Enhanced safety features in 3.6 Flash reduce vulnerability to jailbreaks.
- 3.5 Flash-Lite significantly outperforms previous Flash-Lite generations.
DeepSignal Analysis
What happened
Google DeepMind has launched three new AI models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model reportedly reduces output token usage by 17% compared to its predecessor, while 3.5 Flash-Lite achieves a throughput of 350 tokens per second. These models aim to enhance efficiency and performance for developers building AI agents.
Key evidence
- Gemini 3.6 Flash reduces output token usage by 17% compared to 3.5 Flash, improving efficiency for developers.
- 3.5 Flash-Lite operates at 350 output tokens per second, making it the fastest model in the 3.5 series.
- 3.5 Flash Cyber is designed for cybersecurity applications, working alongside the CodeMender infrastructure to detect and fix vulnerabilities.
Why it matters
The introduction of these models reflects a growing demand for more efficient AI solutions in production environments. By reducing token usage and improving throughput, developers can potentially lower costs and enhance the performance of AI agents. The focus on cybersecurity with the Flash Cyber model also highlights the increasing importance of secure AI applications in today's digital landscape.
Source Excerpt
We’re introducing new Gemini models, including Gemini 3. 6 Flash, 3. 5 Flash-Lite and 3. 5 Flash Cyber.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Google DeepMind
See more →
Introducing Gemma 4 12B: a unified, encoder-free
Google DeepMind has introduced Gemma 4 12B, a unified, encoder-free multimodal model designed to enhance performance across various tasks. This model aims to streamline processes in AI applications by eliminating the need for traditional encoders, potentially improving efficiency and reducing costs for developers and researchers in the field.

