Google releases three new models including Gemini 3.6 Flash: saves 17% on token costs and is twice as fast

robot
Abstract generation in progress

Author: Google DeepMind

Translated by: ShenChao TechFlow

ShenChao takeaway: The core of all three models released by Google this time is addressing the real pain points of commercializing AI Agents—cost and speed. Compared with the previous generation, 3.6 Flash saves 17% of tokens, with a lower price but better quality; 3.5 Flash-Lite reaches 350 tokens/second, making it the fastest 3.5-series model currently available; and the dedicated network security model Flash Cyber targets the urgent need market of fixing code vulnerabilities. This isn’t a performance benchmark game—it’s paving the way for large-scale deployment of AI Agents.

Google DeepMind today released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. These models are designed specifically for the efficiency, latency, and reliability required to build AI Agents at scale.

Gemini 3.6 Flash: higher efficiency, better quality

3.6 Flash is directly built on improvements to developer feedback on 3.5 Flash. It not only goes further in coding and knowledge work, but also significantly improves token efficiency. According to the Artificial Analysis Index, 3.6 Flash reduces output token consumption by 17% compared with 3.5 Flash. In certain benchmarks such as Datacurve’s DeepSWE, the reduction can be as high as 65%. It also requires fewer reasoning steps and tool calls to complete multi-step workflows.

These efficiency gains come with a lower price as well. Pricing is $1.5 per million input tokens and $7.5 per million output tokens. 3.6 Flash lowers the total cost per Agent task, making it more economical to build and run Agents.

Even with higher efficiency, 3.6 Flash performs better than 3.5 Flash across various use cases:

More precise code editing, reducing unnecessary changes and execution loops, reaching 49% on DeepSWE (3.5 Flash: 37%), and improving significantly to 63.9% on the machine learning research benchmark MLE Bench (3.5 Flash: 49.7%)

Improved computer operation capability, with an OSWorld-Verified score of 83.0% (3.5 Flash: 78.4%). Computer operation is now provided as a built-in client tool via the Gemini API and Gemini Enterprise

Better performance in knowledge work, such as a GDPval-AA v2 benchmark score of 1421 (3.5 Flash: 1349). Customers including Hebbia and Harvey have found it especially strong in multimodal tasks like document parsing, chart and data analysis, and report drafting

Customer feedback says 3.6 Flash is an improvement in both cost and quality, balancing token efficiency, accuracy, and speed in complex workflows and knowledge-oriented tasks.

Security design

3.6 Flash is equipped with enhanced frontier safety protections covering chemical, biological, radiological, and nuclear (CBRN) as well as cyber attack abuse. These protections greatly improve the model’s resistance to jailbreak attacks. At the same time, the model is trained to minimize refusals for beneficial uses as much as possible.

Gemini 3.5 Flash-Lite: built for scaling Agent workflows

3.5 Flash-Lite is designed specifically for low-latency tasks and developer scenarios that require high throughput, such as Agent search and document processing.

3.5 Flash-Lite is the fastest model in the 3.5 series. Measured by Artificial Analysis, its speed reaches 350 output tokens per second. Pricing is $0.3 per million input tokens and $2.5 per million output tokens, and its quality is significantly better than 3.1 Flash-Lite, offering excellent value for developers and customers running high-volume production tasks.

3.5 Flash-Lite supports efficient scaling for Agent systems. At all levels of thinking, the model significantly outperforms 3.1 Flash-Lite. Developers can configure the model based on workload: use the minimum and lower thinking levels for large batches of tasks, prioritizing low latency and low-cost execution; or enable higher thinking levels to handle multi-step sub-Agent workloads. The model now also includes computer operation as a built-in tool, reliably supporting Agent tasks across interfaces.

It also sees significant improvements in coding and Agent tasks, such as a 54% score on Terminal-Bench 2.1 (3.1 Flash-Lite: 31%), a 72.2% score on long-context tasks like GDM-MRCR v2 (3.1 Flash-Lite: 60.1%), and a 1140 score in real task execution like GDPval-AA v2 (3.1 Flash-Lite: 642).

In practice, in many Agent and coding evaluations, 3.5 Flash-Lite even surpasses 3 Flash, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%), becoming a faster and stronger choice for 2.5 and 3 Flash workloads.

Early customers emphasized 3.5 Flash-Lite’s unique combination of speed, intelligence, and cost efficiency for scaling Agent workflows and data processing tasks.

Gemini 3.5 Flash Cyber in CodeMender: efficiently discover and fix vulnerabilities

The speed at which AI models discover security vulnerabilities has already surpassed the speed at which current systems fix them. Addressing this increasingly serious threat requires a software security approach that is both efficient and powerful.

Flash’s performance and efficiency make it an ideal foundation for large-scale detection, verification, and patching of code security issues. Gemini 3.5 Flash Cyber is built on 3.5 Flash, micro-tuned specifically to discover and fix network security vulnerabilities, and with a per-token price lower than large models.

In CodeMender, multiple 3.5 Flash Cyber Agents work together to generate a single comprehensive report, reaching cutting-edge competitive levels on the popular CyberGym benchmark.

Given the dual-use nature of this technology, we take a cautious deployment approach. The model will soon be provided specifically through CodeMender in the form of limited-access pilot projects for government and trusted partners. This will allow frontline defenders to discover and fix critical vulnerabilities before they are exploited, while also mitigating broader abuse risks.

Get started immediately

3.6 Flash and 3.5 Flash-Lite are available starting today:

Developers can use the Gemini API through Google AI Studio and Android Studio. 3.6 Flash can also be used in Google Antigravity

Enterprises can use it on the Gemini Enterprise Agent platform. 3.6 Flash can also be used in Gemini Enterprise applications

Everyone can use it via the Gemini app. 3.5 Flash-Lite is also rolling out in Google Search

Share feedback to help improve future Gemini models—we look forward to releasing 3.5 Pro soon.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned