Gemini 3.7 Flash: Google Halves the Price of Its Newest Model Until Year-End
Set Trending Topics as a preferred source on Google.
Just three weeks after Gemini 3.6 Flash, Google released the next model in its Flash line on August 13. Gemini 3.7 Flash targets coding, agentic workflows and document processing in enterprise settings – and arrives with an introductory price that halves what Flash models cost so far.
The price is the real argument
Through December 31, 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens via the API. From January 1, 2027, list prices double to $1.50 and $7.50 respectively. According to Google, the discount also applies retroactively to Gemini 3.6 Flash, which had only launched at the higher prices in late July.
For teams running agents in production, that time window matters more than any single benchmark row: at high token volumes, inference costs add up fast. VentureBeat argues that the months until year-end give companies a chance to test whether Google’s promise of fewer retries and less manual oversight actually translates into lower total operating costs.
What Google claims has improved
Google calls the model its most intelligent “workhorse” model yet for coding and agents. According to the company, it adapts better when it hits roadblocks, asks for clarification more often when a task is ambiguous, and follows instructions more closely. For multi-step planning and tool calls, it is said to invest more compute. Worth highlighting is a very high throughput of 340 tokens per second, the kind of speed real-time applications require.
The benchmark figures Google published, compared with its predecessor 3.6 Flash:
| Benchmark | 3.7 Flash | 3.6 Flash |
|---|---|---|
| FrontierCode 1.1 Main (production code) | 43.6% | 34.4% |
| DeepSWE v1.1 (long-horizon software engineering) | 65.3% | 49.0% |
| WebDev Arena (Elo) | 1,588 | 1,538 |
| GDP.pdf (complex documents) | 34.0% | 22.0% |
| AutomationBench (enterprise workflows) | 30.4% | 17.0% |
Technically, according to the model card, this is not a new pretraining run but a refinement of 3.6 Flash with algorithmic improvements to the reasoning core. The context window is one million tokens, maximum output is 64,000 tokens, and the knowledge cutoff remains March 2026.
No across-the-board lead over the competition
All of these numbers come from Google’s own testing; independent verification is not yet available. And even in Google’s own comparison table, 3.7 Flash does not win on every front: on Terminal-bench 2.1 it lands at 85.8%, behind GPT-5.6 Terra (87.4%), and the OpenAI model also leads on Terminal-bench 3.0 and OSWorld-2.0. On the multimodal Agent’s Last Exam, Claude Sonnet 5 leads with 33.3% against Gemini 3.7 Flash’s 26.3%. Conversely, Google pulls clearly ahead on AutomationBench and GDP.pdf.
Overall, 3.7 Flash sits somewhat behind the current top models from OpenAI, Anthropic and SpaceXAI – and also behind the leading open-weights models from Alibaba (Qwen) and Moonshot AI (KimiK3).
The flagship that is still missing
Developers can access 3.7 Flash through the Gemini API in Google AI Studio and Android Studio, as well as in the agent environment Google Antigravity. For enterprises, it is available in the Gemini Enterprise Agent Platform and the Gemini Enterprise app. End users with a Google AI Pro or Ultra subscription get the model via Gemini Spark, the always-on assistant offered in more than 160 countries.
Google also states that it has shipped updated safeguards against misuse in the CBRN (chemical, biological, radiological, nuclear) and cyber offense domains.
The pace on the Flash line remains striking – as does the gap above it: once again, no release date was given for the more powerful flagship Gemini 3.5 Pro, which has been delayed for months.

