Google releases Gemini 3.6 Flash and 3.5 Flash-Lite updates, intensifying AI price war

Google releases Gemini 3.6 Flash and 3.5 Flash-Lite updates, intensifying AI price war

Google expands its Gemini lineup with faster and cheaper models for coding, agentic workflows and cybersecurity.

Google has launched three new Gemini models aimed at making AI agents faster, cheaper and more efficient as competition intensifies across coding, enterprise automation and cybersecurity.

The lineup includes Gemini 3.6 Flash, Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber. Google also confirmed that Gemini 3.5 Pro is being tested with partners and that it has begun its most ambitious pretraining run yet for Gemini 4.

Gemini 3.6 Flash serves as the main model in the release, offering stronger coding, knowledge work and multimodal performance while consuming fewer tokens than Gemini 3.5 Flash.

According to Google, the model uses 17% fewer output tokens on the Artificial Analysis Index. The company said token reductions reached as much as 65% in certain DeepSWE tests, with the model also requiring fewer reasoning steps and tool calls to complete multistep workflows.

Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens, giving developers a lower price per agentic task than its predecessor.

The model scored 49% on DeepSWE, compared with 37% for Gemini 3.5 Flash, reflecting higher precision and fewer unnecessary code changes. Its MLE Bench score increased to 63.9% from 49.7%.

Computer use also improved, with Gemini 3.6 Flash scoring 83% on OSWorld Verified, compared with 78.4% for the previous model. Computer use is available as a built in client side tool through the Gemini API and Gemini Enterprise.

Google said the model is particularly suited for agentic workflows, coding tasks and long running enterprise processes. It supports text, images, audio and video, with a context window of up to one million tokens.

Advertisement

Flash Lite targets high volume AI workloads

Gemini 3.5 Flash Lite is designed for applications where speed, throughput and cost are more important than using a larger model.

The model generates around 350 output tokens per second, according to Artificial Analysis, making it the fastest model in the Gemini 3.5 lineup.

It is priced at $0.30 per million input tokens and $2.50 per million output tokens. Google is positioning it for high volume workloads such as agentic search, document processing, data extraction, translation and summarization.

Developers can adjust the model’s thinking level depending on the task. Minimal and low settings prioritize faster and cheaper execution, while higher settings support more complex subagent workflows.

Gemini 3.5 Flash Lite scored 54% on Terminal Bench 2.1, compared with 31% for Gemini 3.1 Flash Lite. It also reached 72.2% on GDM MRCR v2, up from 60.1%, and 1,140 on GDPval AA v2, compared with 642.

Google said Flash Lite also outperformed the larger Gemini 3 Flash on several coding and agent benchmarks, including SWE Bench Pro and OSWorld Verified.

Google positions Flash Cyber against Anthropic’s Mythos

The third model, Gemini 3.5 Flash Cyber, is specialized for finding, validating and fixing software vulnerabilities.

Built on Gemini 3.5 Flash, the model operates through CodeMender, Google’s code security agent. Multiple Flash Cyber agents can work together to analyze vulnerabilities and produce a combined report.

Google said the system achieved competitive performance on CyberGym while operating at a lower price per token than larger cybersecurity models.

The release positions Flash Cyber as a lower cost challenger to advanced security models such as Anthropic’s Mythos, which is designed for vulnerability discovery and complex cyber reasoning.

Google is taking a restricted approach to the model’s deployment due to the potential for cyber misuse. Flash Cyber will initially be available only to governments and trusted partners through a limited CodeMender pilot.

The company said the rollout is intended to give defenders access to vulnerability detection and remediation tools before the model is made more broadly available.

Gemini 4 pretraining begins

Gemini 3.6 Flash and Gemini 3.5 Flash Lite are available through the Gemini API, Google AI Studio, Android Studio and Gemini Enterprise.

Gemini 3.6 Flash is also available through Google Antigravity and the Gemini app, while Flash Lite is rolling out through Google Search.

The releases arrive while Google continues testing Gemini 3.5 Pro with partners. The flagship model has not yet received a broad release date.

Google also confirmed that development of its next generation is underway, with the company beginning what it described as its most ambitious pretraining run to date for Gemini 4.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

Google releases Gemini 3.6 Flash and 3.5 Flash-Lite updates, intensifying AI price war

Google releases Gemini 3.6 Flash and 3.5 Flash-Lite updates, intensifying AI price war

Google expands its Gemini lineup with faster and cheaper models for coding, agentic workflows and cybersecurity.

Share

Add us on Google

Google has launched three new Gemini models aimed at making AI agents faster, cheaper and more efficient as competition intensifies across coding, enterprise automation and cybersecurity.

The lineup includes Gemini 3.6 Flash, Gemini 3.5 Flash Lite and Gemini 3.5 Flash Cyber. Google also confirmed that Gemini 3.5 Pro is being tested with partners and that it has begun its most ambitious pretraining run yet for Gemini 4.

Gemini 3.6 Flash serves as the main model in the release, offering stronger coding, knowledge work and multimodal performance while consuming fewer tokens than Gemini 3.5 Flash.

According to Google, the model uses 17% fewer output tokens on the Artificial Analysis Index. The company said token reductions reached as much as 65% in certain DeepSWE tests, with the model also requiring fewer reasoning steps and tool calls to complete multistep workflows.

Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output tokens, giving developers a lower price per agentic task than its predecessor.

The model scored 49% on DeepSWE, compared with 37% for Gemini 3.5 Flash, reflecting higher precision and fewer unnecessary code changes. Its MLE Bench score increased to 63.9% from 49.7%.

Computer use also improved, with Gemini 3.6 Flash scoring 83% on OSWorld Verified, compared with 78.4% for the previous model. Computer use is available as a built in client side tool through the Gemini API and Gemini Enterprise.

Google said the model is particularly suited for agentic workflows, coding tasks and long running enterprise processes. It supports text, images, audio and video, with a context window of up to one million tokens.

Advertisement

Flash Lite targets high volume AI workloads

Gemini 3.5 Flash Lite is designed for applications where speed, throughput and cost are more important than using a larger model.

The model generates around 350 output tokens per second, according to Artificial Analysis, making it the fastest model in the Gemini 3.5 lineup.

It is priced at $0.30 per million input tokens and $2.50 per million output tokens. Google is positioning it for high volume workloads such as agentic search, document processing, data extraction, translation and summarization.

Developers can adjust the model’s thinking level depending on the task. Minimal and low settings prioritize faster and cheaper execution, while higher settings support more complex subagent workflows.

Gemini 3.5 Flash Lite scored 54% on Terminal Bench 2.1, compared with 31% for Gemini 3.1 Flash Lite. It also reached 72.2% on GDM MRCR v2, up from 60.1%, and 1,140 on GDPval AA v2, compared with 642.

Google said Flash Lite also outperformed the larger Gemini 3 Flash on several coding and agent benchmarks, including SWE Bench Pro and OSWorld Verified.

Google positions Flash Cyber against Anthropic’s Mythos

The third model, Gemini 3.5 Flash Cyber, is specialized for finding, validating and fixing software vulnerabilities.

Built on Gemini 3.5 Flash, the model operates through CodeMender, Google’s code security agent. Multiple Flash Cyber agents can work together to analyze vulnerabilities and produce a combined report.

Google said the system achieved competitive performance on CyberGym while operating at a lower price per token than larger cybersecurity models.

The release positions Flash Cyber as a lower cost challenger to advanced security models such as Anthropic’s Mythos, which is designed for vulnerability discovery and complex cyber reasoning.

Google is taking a restricted approach to the model’s deployment due to the potential for cyber misuse. Flash Cyber will initially be available only to governments and trusted partners through a limited CodeMender pilot.

The company said the rollout is intended to give defenders access to vulnerability detection and remediation tools before the model is made more broadly available.

Gemini 4 pretraining begins

Gemini 3.6 Flash and Gemini 3.5 Flash Lite are available through the Gemini API, Google AI Studio, Android Studio and Gemini Enterprise.

Gemini 3.6 Flash is also available through Google Antigravity and the Gemini app, while Flash Lite is rolling out through Google Search.

The releases arrive while Google continues testing Gemini 3.5 Pro with partners. The flagship model has not yet received a broad release date.

Google also confirmed that development of its next generation is underway, with the company beginning what it described as its most ambitious pretraining run to date for Gemini 4.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.