August 14, 2026AgentsAPICoding

Gemini 3.7 Flash: three weeks later, half the price

Google shipped Gemini 3.7 Flash three weeks after 3.6 Flash. Three weeks. And the introductory price through end of 2026 is $0.75 per million input and $3.75 per million output, which is half what 3.6 Flash cost at launch. When a company reprices a model to half within a month of the last one, that's not a roadmap, that's a correction.

The numbers explain the hurry. DeepSWE v1.1 goes 49.0 to 65.3 percent. FrontierCode 1.1 Main goes 34.4 to 43.6. WebDev Arena Elo moves 1538 to 1588. Document handling on GDP.pdf jumps 22.0 to 34.0 percent, which for anyone running agents over invoices and contracts is the line that actually changes a workflow. Google calls it their most intelligent workhorse model yet for coding and agents, and for once the benchmark spread supports the adjective.

The qualitative claims matter more than the numbers for agent builders. Google says the model better adapts to roadblocks, asks for clarification when intent is genuinely ambiguous, and follows instructions with higher fidelity. Those three behaviors are the difference between an agent that burns forty turns confidently doing the wrong thing and one that stops to ask. Instruction fidelity is the boring metric that decides whether your harness works.

Context worth holding: Gemini 3.5 Pro is still delayed. The flagship has slipped repeatedly while the Flash line ships on a three-week cadence and keeps cutting price. Google is winning the workhorse tier by volume and cadence while the frontier tier waits, which is a coherent strategy if you believe most agent tokens are cheap tokens. Most agent tokens are cheap tokens.

Live now in AI Studio, Vertex, Android Studio, and rolling out in Spark inside the Gemini app for AI Pro and Ultra subscribers.
← Previous
DeepSeek Harness: everything is a plugin, and that's the whole point
Next β†’
750 tokens a second changes what an agent is for
← Back to all articles

Comments

Loading...
>_