model

Google Launches Gemini 3.6 Flash: Fast, Affordable, Safe

Google introduces token-efficient Gemini 3.6 Flash, plus 3.5 Flash-Lite and 3.5 Flash Cyber for efficiency and security.

18:08 UTC · Jul 223 min readLintasAI Editorial Desk
Google Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber models illustration.

Google DeepMind officially launched three new models in the Gemini Flash lineup: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. These models are designed to meet the needs of developers building production-scale AI agents, with a focus on higher token efficiency, low latency, and more reliable performance.

What's New with Gemini 3.6 Flash?

Gemini 3.6 Flash arrives as the "workhorse" model, succeeding 3.5 Flash, with leaps in coding performance, knowledge work, and multimodal capabilities. According to the Artificial Analysis Index, this model uses 17% fewer output tokens than its predecessor. On the DeepSWE benchmark, token savings reach up to 65%.

Efficiency does not sacrifice quality. DeepSWE score rose from 37% (3.5 Flash) to 49% (3.6 Flash). On MLE Bench, ML research capability jumped from 49.7% to 63.9%. Computer use on OSWorld-Verified reached 83.0%, up from 78.4%. Knowledge measured by GDPval-AA v2 increased from 1349 to 1421.

Pricing is also lower: US$1.50 per 1M input tokens and US$7.50 per 1M output tokens. "Customers report that 3.6 Flash is a step forward in cost and quality, balancing token efficiency, accuracy, and speed across complex workflows," wrote Google DeepMind.

On the safety side, 3.6 Flash includes stricter Frontier Safety protections against CBRN and cyber offense misuse, and is more resilient to jailbreaking without reducing utility for legitimate applications.

Gemini 3.5 Flash-Lite: Fastest and Most Cost-Efficient Model?

Designed for high-volume, low-latency tasks, Gemini 3.5 Flash-Lite offers a speed of 350 output tokens per second (per Artificial Analysis). At US$0.30 per 1M input tokens and US$2.50 per 1M output tokens, this model has an aggressive price-to-performance ratio.

Compared to 3.1 Flash-Lite, improvements are significant: Terminal-Bench 2.1 score (54% vs 31%), GDM-MRCR v2 for long context (72.2% vs 60.1%), and GDPval-AA v2 (1140 vs 642). In several agent and coding tests, 3.5 Flash-Lite outperformed Gemini 3 Flash, such as on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%).

The model supports configurable thinking levels: minimal or low for latency and cost priorities, or high for multi-step sub-agent tasks. Computer use is also included as a built-in tool.

3.5 Flash Cyber and CodeMender for Cybersecurity

A security-specific version, Gemini 3.5 Flash Cyber, is a model fine-tuned to detect, validate, and fix code vulnerabilities. It is integrated into CodeMender, a code security agent that uses multiple 3.5 Flash Cyber agents to generate a consolidated report.

On the CyberGym benchmark, this combination shows competitive frontier performance. Due to its dual-use nature, the model will only be available through a restricted access program for government and trusted partners. "This gives defenders an early opportunity to find and fix critical vulnerabilities before they are exploited," explained Google DeepMind.

What This Means for Developers

For developers and companies, these models open opportunities to build more efficient and affordable AI agents via the Gemini API in Google AI Studio, Android Studio, and Google Antigravity. Built-in multimodal and computer use capabilities simplify task automation such as document analysis, data processing, or interface prototyping. 3.5 Flash-Lite is already rolling out in Google Search. While waiting for the upcoming Gemini 3.5 Pro, developers can already experiment and integrate these Flash models to reduce operational costs while boosting performance.

Related briefs