Model Release

Google Launches Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on July 21, 2026. Prices, benchmarks, and what they mean for AI agent builders.

LUMIEN5 min read
Google Launches Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google DeepMind released three new Gemini models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. The headline number is a 17% reduction in output token usage for 3.6 Flash versus its predecessor, paired with a lower price per token. Flash-Lite tops out at 350 output tokens per second, making it the fastest in the 3.5 series. The release targets developers building production AI agents who need lower costs, less latency, and more predictable throughput at scale.

What happened

Detail Value
Release date July 21, 2026
3.6 Flash input price $1.50 per 1M tokens
3.6 Flash output price $7.50 per 1M tokens
3.5 Flash-Lite input price $0.30 per 1M tokens
3.5 Flash-Lite output price $2.50 per 1M tokens
3.5 Flash-Lite speed (Artificial Analysis) 350 output tokens/second
3.6 Flash token reduction vs 3.5 Flash 17% fewer output tokens (up to 65% on DeepSWE)

Google DeepMind published announcements for all three models simultaneously, framing them as purpose-built for teams scaling agentic workflows, where token count and latency directly affect operating costs.

Gemini 3.6 Flash: cheaper and sharper than 3.5 Flash

Gemini 3.6 Flash is positioned as the workhorse upgrade. According to the Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash on average, and on the DeepSWE benchmark by Datacurve the reduction reaches up to 65%. Fewer tokens means lower bills for any workflow running at volume.

Performance also moves forward on several tested benchmarks compared to 3.5 Flash:

  • DeepSWE (coding): 49% vs. 37%
  • MLE Bench (ML research): 63.9% vs. 49.7%
  • OSWorld-Verified (computer use): 83.0% vs. 78.4%
  • GDPval-AA v2 (knowledge work): 1421 vs. 1349

Computer use is now a built-in client-side tool via the Gemini API and Gemini Enterprise, rather than something developers need to wire up separately. Customers named in the announcement include Hebbia and Harvey, both of which use 3.6 Flash for document parsing, chart analysis, and report drafting.

On safety, Google says 3.6 Flash ships with reinforced safeguards against misuse in Chemical, Biological, Radiological, and Nuclear (CBRN) domains and cyber offense. The model card notes it has been trained to keep refusals low for legitimate use cases while being more resistant to jailbreaks.

Gemini 3.5 Flash-Lite: built for throughput

Flash-Lite is the speed tier. At 350 output tokens per second (Artificial Analysis), it is the fastest model in the 3.5 series. Its pricing is significantly lower than 3.6 Flash: $0.30 per 1M input tokens and $2.50 per 1M output tokens.

Compared to its predecessor, 3.1 Flash-Lite, the new version shows better quality across all thinking levels. Developers can dial the model between minimal, low, and higher thinking levels depending on whether a task needs raw speed or multi-step reasoning for sub-agent workloads. Computer use is also a built-in tool here, consistent with 3.6 Flash.

Gemini 3.5 Flash Cyber in CodeMender

The third release is more specialized. Gemini 3.5 Flash Cyber is a cybersecurity-focused model paired with Google’s CodeMender code security agent. Google describes this as a combination of a specialized model and an agent infrastructure layer, rather than a general-purpose release. The announcement claims competitive frontier performance on cybersecurity tasks, though specific benchmark numbers for this model were not published.

What else Google signaled

Google DeepMind confirmed that Gemini 3.5 Pro is currently in partner testing, with a broad release coming once it is ready. The team also disclosed that pre-training has started for Gemini 4, which Google describes as its most ambitious pre-training run yet.

Why it matters

Token efficiency is the real story here. For any business running AI agents, output token count is one of the largest variables in cost. A 17% average reduction, and up to 65% on coding tasks, can meaningfully change the economics of a production pipeline. If you are currently on 3.5 Flash, the price is also lower on 3.6 Flash, so there is no cost penalty for upgrading.

Flash-Lite at $0.30 input and $2.50 output is a credible choice for high-volume classification, summarization, or document processing where you need throughput rather than deep reasoning. The adjustable thinking levels make it more flexible than a single-tier speed model.

Teams exploring AI integration for document-heavy workflows, such as invoice processing or contract review, now have a clearer cost structure to model before committing to a stack.

Our take

The efficiency angle is more interesting than the benchmark jump. Most teams we talk to are not constrained by model quality at this point; they are constrained by the cost of running agents at volume. A 17% token reduction that compounds across every agent call is the kind of improvement that actually changes a business case, not a headline score on a leaderboard.

The Flash-Lite tier is worth a closer look for anyone doing agentic search or batch document processing. At 350 tokens per second and $2.50 per 1M output tokens, the math gets interesting fast for high-volume tasks that do not need deep reasoning.

The CodeMender and Flash Cyber release is the most opaque. “Competitive at the frontier” without published numbers is a soft claim. We would treat it as early access signal rather than a proven production choice until more benchmark data surfaces.

With Gemini 3.5 Pro in partner testing and Gemini 4 pre-training already running, the cadence is accelerating. If you are building on Gemini now, plan for another migration decision within a few months. Our coverage of the best local LLMs for constrained environments is worth reading alongside this if cost control is your primary concern.

What to do about it

  1. Run a token count comparison on your current 3.5 Flash workload for one week, then estimate savings at 3.6 Flash pricing before switching.
  2. Test Flash-Lite on any batch pipeline where you are prioritizing throughput over reasoning depth, using the minimal thinking level first.
  3. If you use 3.5 Flash for coding or document tasks, benchmark 3.6 Flash on your actual inputs against the DeepSWE and GDPval results, since your workload may differ from lab conditions.
  4. Hold on CodeMender and Flash Cyber until third-party benchmark results appear.
  5. If you want help modeling the cost and quality tradeoffs for your specific agent setup, get in touch with the Lumien team.

The short version: upgrade 3.5 Flash users to 3.6 Flash now, pilot Flash-Lite for bulk tasks, and watch for Gemini 3.5 Pro benchmarks before making any larger architectural decisions.

Source: Google DeepMind

Frequently asked questions

How much does Gemini 3.6 Flash cost?

Gemini 3.6 Flash is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, which is lower than Gemini 3.5 Flash.

How fast is Gemini 3.5 Flash-Lite?

According to Artificial Analysis, Gemini 3.5 Flash-Lite runs at 350 output tokens per second, making it the fastest model in the 3.5 series.

How does Gemini 3.6 Flash compare to 3.5 Flash on benchmarks?

3.6 Flash scores 49% vs 37% on DeepSWE, 63.9% vs 49.7% on MLE Bench, and 83.0% vs 78.4% on OSWorld-Verified, while using 17% fewer output tokens on average.

When is Gemini 4 coming out?

Google DeepMind has confirmed that pre-training for Gemini 4 has started, describing it as their most ambitious pre-training run yet, but no release date has been announced.

More from AI