Gemini 4 Argon Ships to Cyber Defenders First, Everyone Else Waits
Google's most capable model is out, and almost nobody can use it. Gemini 4 Argon launched on September 30, but it is going only to trusted cyber defenders in Google's Fairwind Program. Google says it will widen access "as soon as possible," after more guardrail work and the US government's voluntary pre-release review. That is now the standard script: Anthropic did it with Mythos, and Google is doing it here.
The benchmarks are built for agents. Argon sets a new high on DeepSWE v1.1 (77.9%), a long-horizon software-engineering test. It ranks first on Zapier's AutomationBench at 51.3%, leads the Vals Index (finance, coding, legal and tax work, weighted by GDP share), and scores 91.7% on LVBench for long video. It ties for first on CWE-bench v1 at 68% for fixing vulnerabilities. The biggest spec change: the output limit goes from 64K tokens to 1 million. Google's argument is that a model given room to generate hundreds of thousands of tokens in one trajectory can finish hard problems in one go.
The internal case studies beat the leaderboards. A team of Argon agents read fleet-wide profiling data and freed more than 300 TiB of memory across Google's data centers, with 500 TiB to 1 PiB expected in total. Argon agents are porting C and C++ to Rust, up to the 800K-line Fuchsia Zircon kernel. In libgav1, Google's video decoder, Argon took an existing Rust port and replaced 32K lines of hand-written SIMD with safe Rust that the compiler vectorizes on its own, through rounds of profile-guided experiments. The result runs 2.7x faster than that port, with identical output. That is agent work, not chat.
The cyber angle cuts both ways. Trusted defenders get Argon with no cyber guardrails at all. Wiz's Scan for Good program has already used it to find a critical flaw exposing patient data in hospital software worldwide, one previous frontier models had missed. For everyone else, Google is adding activation-level misuse monitoring and claims the lead on Gray Swan's indirect prompt-injection benchmark.
One more number to notice: introductory pricing is $2 per million input tokens and $10 per million output, with cached input 95% off. That is exactly Sonnet 5.5's price, for what Google calls its flagship. If that price holds at general availability, the frontier just got a lot cheaper, but nobody outside Fairwind can test it yet.
Link: blog.google (Gemini 4 Argon)
← Back to all articles
The benchmarks are built for agents. Argon sets a new high on DeepSWE v1.1 (77.9%), a long-horizon software-engineering test. It ranks first on Zapier's AutomationBench at 51.3%, leads the Vals Index (finance, coding, legal and tax work, weighted by GDP share), and scores 91.7% on LVBench for long video. It ties for first on CWE-bench v1 at 68% for fixing vulnerabilities. The biggest spec change: the output limit goes from 64K tokens to 1 million. Google's argument is that a model given room to generate hundreds of thousands of tokens in one trajectory can finish hard problems in one go.
The internal case studies beat the leaderboards. A team of Argon agents read fleet-wide profiling data and freed more than 300 TiB of memory across Google's data centers, with 500 TiB to 1 PiB expected in total. Argon agents are porting C and C++ to Rust, up to the 800K-line Fuchsia Zircon kernel. In libgav1, Google's video decoder, Argon took an existing Rust port and replaced 32K lines of hand-written SIMD with safe Rust that the compiler vectorizes on its own, through rounds of profile-guided experiments. The result runs 2.7x faster than that port, with identical output. That is agent work, not chat.
The cyber angle cuts both ways. Trusted defenders get Argon with no cyber guardrails at all. Wiz's Scan for Good program has already used it to find a critical flaw exposing patient data in hospital software worldwide, one previous frontier models had missed. For everyone else, Google is adding activation-level misuse monitoring and claims the lead on Gray Swan's indirect prompt-injection benchmark.
One more number to notice: introductory pricing is $2 per million input tokens and $10 per million output, with cached input 95% off. That is exactly Sonnet 5.5's price, for what Google calls its flagship. If that price holds at general availability, the frontier just got a lot cheaper, but nobody outside Fairwind can test it yet.
Link: blog.google (Gemini 4 Argon)
Comments