July 22, 2026AgentsAPIBenchmark

Gemini 3.6 Flash Is Google Tuning for Agents, Not Chat

Google shipped three models on July 21 and skipped the one everyone was waiting for. No Gemini 3.5 Pro. What landed instead was Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a specialized security model called 3.5 Flash Cyber. Read the release notes and the priority is obvious: this is a lineup built for things that run in a loop, not things that answer a question.

The headline number on 3.6 Flash is not a capability score, it is token consumption. It uses 17% fewer output tokens than 3.5 Flash on average, up to 65% fewer on some benchmarks. That is an agent metric. Nobody chatting with a model cares about output token count. Somebody running a hundred-step trajectory cares about nothing else. Pricing is $1.50 per million input and $7.50 per million output. Capability moved too: DeepSWE 49% against 37%, MLE Bench 63.9% against 49.7%, OSWorld-Verified 83.0% against 78.4%. Computer use is built in, not bolted on.

Flash-Lite is the cheap fast one and it is the more interesting release. $0.30 in, $2.50 out, 350 output tokens a second, configurable thinking levels, computer use also built in. Terminal-Bench 2.1 went from 31% to 54%. Long context tasks 60.1% to 72.2%. On SWE-Bench Pro it scores 54.2% and beats Gemini 3 Flash at 49.6%, which means the cheapest model in the new lineup outperforms a flagship from one generation back on real coding work. That is the actual shape of the curve right now.

Then there is Flash Cyber, fine-tuned for vulnerability detection and patching, running inside CodeMender through multi-agent orchestration, competitive with frontier models on CyberGym. Google is not putting it in the API. It is a limited-access pilot for governments and trusted partners only. That decision landed the same week OpenAI published a report on its own models breaking containment during a cyber benchmark, which is either good timing or no coincidence at all.

Both 3.6 Flash and 3.5 Flash-Lite are live now in Google AI Studio, Android Studio, Gemini Enterprise and the Gemini app, with Flash-Lite rolling into Search. The full announcement is on blog.google. The read here is that Google has decided the volume business is agents burning tokens by the billion, and it is pricing and tuning accordingly.
← Previous
OpenAI's Own Models Escaped the Sandbox and Hacked Hugging Face
Next β†’
Jack Dorsey Built a Slack Where the Agents Carry Passports
← Back to all articles

Comments

Loading...
>_