Gemini 3.7 Flash: Better Coding Benchmarks at Half the Price of 3.6
Google released Gemini 3.7 Flash on Aug 13, 2026. Key facts: $0.75/1M input tokens, FrontierCode 43.6%, DeepSWE 65.3%, 1M context window, API only.

Google released Gemini 3.7 Flash on August 13, 2026, three weeks after Gemini 3.6 Flash. It is not a new base model but an algorithmic refinement of 3.6 Flash with a stronger focus on coding, document work, and enterprise automation. The headline number is price: $0.75 per 1M input tokens and $3.75 per 1M output tokens until December 31, 2026, roughly a third the blended cost of Claude Sonnet 5 or GPT-5.6 Terra at similar input/output ratios.
What happened
| Detail | Value |
|---|---|
| Release date | August 13, 2026 |
| Input price (until Dec 31, 2026) | $0.75 per 1M tokens |
| Output price (until Dec 31, 2026) | $3.75 per 1M tokens |
| Input / output price from Jan 1, 2027 | $1.50 / $7.50 per 1M tokens |
| Context window | 1M tokens |
| Max output tokens | 64K |
| Knowledge cutoff | March 2026 |
| Access | API and enterprise only. No open weights. |
Google describes 3.7 Flash as a refinement of 3.6 Flash, not a new pretraining run. The improvements come from algorithmic changes to the core reasoning layer. The model accepts text, images, audio, and video, and it supports configurable “thinking” modes that let you trade reasoning depth against cost and latency.
Access is available through the Gemini API, Google AI Studio, Google Antigravity, Android Studio, the Gemini Enterprise Agent Platform, and the Gemini Enterprise app. Consumer users can reach it via Gemini Spark on Google AI Pro and Ultra plans. There is no self-hosted or air-gapped deployment option.
How does Gemini 3.7 Flash compare on benchmarks?
| Benchmark | Gemini 3.6 Flash | Gemini 3.7 Flash | GPT-5.6 Terra | Claude Sonnet 5 |
|---|---|---|---|---|
| FrontierCode 1.1 Main | 34.4% | 43.6% | , | , |
| DeepSWE v1.1 | 48.6% | 65.3% | 69.6% | , |
| WebDev Arena (Elo) | 1538 | 1588 | , | , |
| GDP.pdf | 22.0% | 34.0% | , | , |
| AutomationBench | 17.0% | 30.4% | 23.6% | 10.7% |
| GDM-MRCR v2 (128k) | , | 97.0% | , | , |
| GDPval-AA v2 (Elo) | , | 1525 | , | 1598 |
| CharXiv Reasoning | 85.2% | 84.5% | , | , |
| AI Intelligence Index | , | 56 | 57 | , |
The coding and automation gains are the most significant. AutomationBench (a private enterprise workflow evaluation) jumps from 17.0% to 30.4%, which puts 3.7 Flash ahead of both GPT-5.6 Terra (23.6%) and Claude Sonnet 5 (10.7%) on that specific test.
GPT-5.6 Terra still leads on terminal-based and computer-use tasks: Terminal-bench 2.1 (87.4%), Terminal-bench 3.0 (20.8%), OSWorld-2.0 (50.2%), and DeepSWE (69.6% vs 65.3%). CharXiv Reasoning is a small step backward from 3.6 Flash. Neither model is a clean winner across every category.
Why it matters
Price is the real story. At an 80/20 input-to-output token ratio, the blended cost works out to roughly $1.35 per 1M tokens today, compared to $3.60 for Claude Sonnet 5 and $4.00 for GPT-5.6 Terra, according to Google’s own comparison table. That gap matters most for teams running agents continuously, where token costs stack up fast.
Google says the model is best suited for legal, financial services, biosciences, and enterprise operations workflows, citing its Harvey LAB-AA, GDP.pdf, and AutomationBench results as evidence. Long-running coding agents, PDF-to-structured-data pipelines, and UI generation from screenshots are the specific use cases named.
The introductory price expires on December 31, 2026. From January 1, 2027 the rates double to $1.50 input and $7.50 output per 1M tokens. Teams that build workflows around the current price need to account for that in their cost models now.
Our take
Three weeks between model versions is fast, and calling this a “refinement” rather than a new model is honest. The benchmark jumps are real, especially on AutomationBench and FrontierCode, but the comparison table conveniently excludes the categories where GPT-5.6 Terra still wins.
The price argument holds up better than the benchmark argument. For teams doing workflow automation with AI agents at volume, cutting token costs by two-thirds while maintaining competitive reasoning quality is a legitimate reason to run an evaluation. The catch is that you are entirely dependent on Google’s hosted infrastructure. Regulated industries with data-residency requirements are still locked out.
The December 2026 price doubling is worth flagging to any client building a budget around this model today. Lock in the evaluation now and plan your cost model for post-January pricing before committing to any architecture. If you are exploring how AI models like this slot into your existing tools, our AI integration service covers exactly that ground.
What to do about it
- Run your current agent or coding workflow against 3.7 Flash via the Gemini API and compare output quality directly against your existing model.
- Calculate your blended token cost at the actual input/output ratio you use, not a hypothetical 80/20 split.
- Build a cost model for both the introductory rate (until Dec 31, 2026) and the standard rate from January 2027 before committing to any production deployment.
- If your use case involves terminal agents or computer-use tasks, test GPT-5.6 Terra in parallel. 3.7 Flash trails on those benchmarks.
- If you have data-residency requirements, verify your jurisdiction is covered by Google’s enterprise offering before investing in integration work.
The model earns a serious look for document-heavy automation and coding agents. The introductory price window is finite, so evaluate it now rather than after it expires.
Frequently asked questions
How much does Gemini 3.7 Flash cost?
Until December 31, 2026, Gemini 3.7 Flash is priced at $0.75 per 1M input tokens and $3.75 per 1M output tokens. From January 1, 2027 those rates double to $1.50 and $7.50 per 1M tokens respectively.
Is Gemini 3.7 Flash better than Claude Sonnet 5?
It depends on the task. Gemini 3.7 Flash scores higher on AutomationBench (30.4% vs 10.7%) and WebDev Arena, but Claude Sonnet 5 leads on GDPval-AA knowledge work (1598 Elo vs 1525). At current prices, Gemini 3.7 Flash is roughly 2.5x cheaper at a typical token mix.
Can I self-host or run Gemini 3.7 Flash locally?
No. Gemini 3.7 Flash has no open weights and is available only through Google's hosted surfaces, including the Gemini API, Google AI Studio, and enterprise platforms. There is no self-hosted or air-gapped deployment option.
What is the context window size for Gemini 3.7 Flash?
Gemini 3.7 Flash supports a 1 million token context window and returns up to 64,000 output tokens per request. The knowledge cutoff is March 2026.


