August 13, 2026AgentsCodingBenchmark

Grok 4.6 caught GPT-5.6 Sol, and it did it by training on agent tasks

xAI shipped Grok 4.6 yesterday and it scores 61 on the Artificial Analysis Intelligence Index, dead level with GPT-5.6 Sol. CursorBench 3.2 at 69.9%, DeepSWE 1.1 at 65.9%, better than Grok 4.5 across essentially everything. Two dollars per million input, six per million output, with a fast variant at double. Two weeks after 4.5, which is a cadence that says something about how xAI is running.

The pitch is long-running agents. xAI says training included agentic RL tasks spanning knowledge work, coding, and specific environments β€” kernel optimization, web development, CAD. Not synthetic reasoning puzzles. Actual multi-step task environments with real tools and real failure modes. The claimed result is a model that sustains work across many steps, tests its own output, and gets closer on the first try for visual and interactive builds.

That last bit is the underrated one. First-try quality on visual work is not a benchmark you can game, and it is the single thing that decides whether a coding agent feels magical or feels like babysitting. Every retry is context burned, tokens spent, and a human pulled back into the loop.

The distribution is aggressive. Grok 4.6 is live in Cursor, in Grok Build, on the SpaceXAI API, and through OpenRouter, Vercel and Cloudflare, with 2x included usage for the first week. xAI is not waiting for anyone to integrate it. Worth remembering that Artificial Analysis parity is not the same as parity in a hundred-turn agent loop where errors compound β€” but on the leaderboard, the gap to the frontier is now zero.

Details at https://x.ai/news/grok-4-6
← Previous
Qwen just open-weighted a 2.4-trillion-parameter model. Nobody else has done this.
Next β†’
Zed's Delta bets that git commits are the wrong unit for agent work
← Back to all articles

Comments

Loading...
>_