Google launches three new models including Gemini 3.6 Flash! 3.5 Flash-Lite hits breakneck speed, and pre-training for Gemini 4 has already started

According to an announcement on Google’s official blog, Google officially launched a new-generation Flash series model on the 21st, including Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The new models are designed specifically for AI agent workflows, not only significantly lowering output costs and token consumption, but also featuring powerful computer operation capabilities, fully targeting large-scale enterprise production deployments.
(Background: Google reportedly unveils a new “frozen” chip performance surge by 10x! Alphabet’s market cap jumps by $115 billion)
(Background update: Gemini 3.5 Pro has been delayed for months—internal political faction infighting at Google leaves employees frustrated)

Table of contents

Toggle

  • Gemini 3.6 Flash: a major leap in performance, with lower output costs
  • 3.5 Flash-Lite: extremely lightweight, blasting out 350 tokens/s at high speed
  • Focused on network security, Gemini 4 pretraining has started

As artificial intelligence (AI) moves into the agentic era, the balance among large language models in speed, cost, and reasoning capability has become a key consideration for enterprise adoption. On July 21, 2026 (Taipei time), Tulsee Doshi, Senior Product Management Director at Google, published an official article announcing the launch of three new Flash series models built based on prior user feedback, further strengthening its infrastructure advantages in the AI agent space.

Gemini 3.6 Flash: a major leap in performance, with lower output costs

As the new generation’s core working model, Gemini 3.6 Flash has seen notable improvements in code writing, knowledge work, and multimodal task performance. Data shows the model achieved a high score of 83.0% in computer usage capability evaluations (OSWorld-Verified), and the related computer usage tools are now directly built into Gemini API and the enterprise edition.

In addition, 3.6 Flash has significantly optimized token usage efficiency. According to the Artificial Analysis Index, its output token usage is 17% lower than the prior generation, and in some benchmark tests (such as DeepSWE) it can be reduced by as much as 65%. This not only cuts down reasoning and tool-call time in multi-step workflows for AI agents, but also makes its pricing more competitive—input is $1.50 per one million tokens, and output is $7.50.

3.5 Flash-Lite: extremely lightweight, blasting out 350 tokens/s at high speed

For tasks that require low latency and high throughput (e.g., agent search and document processing), Google has introduced a dedicated Gemini 3.5 Flash-Lite. The model’s output speed reaches up to 350 tokens per second, and its input cost is extremely low—just $0.3 per one million tokens (output is $2.5).

Although positioned as lightweight, the quality of the 3.5 Flash-Lite far exceeds the previous Lite generation. It supports adjustable “thinking levels,” allowing developers to freely switch between low-thinking mode and high-thinking mode for handling complex sub-agent workflows. In long-context and agent coding evaluations, it even outperforms the standard Gemini 3 Flash model.

Focused on network security, Gemini 4 pretraining has started

In this update, Google also released a “Gemini 3.5 Flash Cyber” model, which is fine-tuned from 3.5 Flash, for specific cybersecurity defense scenarios. The model can work together with CodeMender, a multi-agent system, to generate comprehensive reports, helping defenders quickly discover and fix network security vulnerabilities. However, considering the dual-use risk of this kind of technology, access is currently limited to limited availability via pilot programs for government institutions and trusted partner organizations only.

At the end of the article, Doshi also hinted that besides the ongoing partner testing for the widely planned Gemini 3.5 Pro, Google has officially started Gemini 4 pretraining as well, describing it as the company’s “most ambitious” plan. Starting today, developers can access the brand-new 3.6 Flash and 3.5 Flash-Lite models via channels such as the Gemini API, getting early experience of the powerful potential of AI agent workflows.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned