Codex Went 18-0 at StarCraft and Still Plays Like a Beginner
Somebody built a version of StarCraft Brood War that can only be played through an agent, originally just to mess around with friends, and it turned into one of the more honest agent benchmarks on the internet. It lives at bw.swerdlow.dev and it does something most benchmarks refuse to do: it tells you the dollar cost per attempt.
Codex Astra won every game at the highest difficulty. Eighteen straight. Average cost ten dollars and fifty-four cents a match. Grok landed between 5.6 and 11.1 percent. Claude Haiku went zero for sixteen. Claude Fable lost a lot but lost interestingly, actually trying to grow an economy and climb a tech tree instead of settling for whatever unit it could afford right now.
The behavioral notes are better than the scores. Codex's strongest recurring idea was disruption, sending a single Probe across the map to bother workers and buildings. It worked absurdly well, mostly because the opposing agent would burn enormous amounts of time deciding what to do about one worker. Grok 4.6's failure is the one every agent engineer will recognize: long stretches of reasoning and very few command batches. It thought beautifully and issued almost no orders, in a game where the clock does not stop while you think. That is not a reasoning deficit, that is a harness and pacing deficit, and it shows up as a loss on the scoreboard either way.
The author's own conclusion is the honest one, and it is why this is worth reading rather than cheering. No agent played beyond beginner level. None could put together a real army composition, none could defend a coordinated attack, none executed anything a human would call a strategy. Codex went 18-0 against the built-in AI and any experienced player would take it apart in ten minutes.
So hold both facts at once. A general-purpose coding agent, with no StarCraft training, beat a game's built-in bot eighteen times in a row at maximum difficulty, which nothing could do three years ago. And it did so while playing badly, at ten dollars a game, in a domain where the skill ceiling is six orders of magnitude above where it landed. Real-time strategy is the cleanest test of the thing agents are worst at, which is acting under a clock, and now we have a price tag for how badly they do it.
← Back to all articles
Codex Astra won every game at the highest difficulty. Eighteen straight. Average cost ten dollars and fifty-four cents a match. Grok landed between 5.6 and 11.1 percent. Claude Haiku went zero for sixteen. Claude Fable lost a lot but lost interestingly, actually trying to grow an economy and climb a tech tree instead of settling for whatever unit it could afford right now.
The behavioral notes are better than the scores. Codex's strongest recurring idea was disruption, sending a single Probe across the map to bother workers and buildings. It worked absurdly well, mostly because the opposing agent would burn enormous amounts of time deciding what to do about one worker. Grok 4.6's failure is the one every agent engineer will recognize: long stretches of reasoning and very few command batches. It thought beautifully and issued almost no orders, in a game where the clock does not stop while you think. That is not a reasoning deficit, that is a harness and pacing deficit, and it shows up as a loss on the scoreboard either way.
The author's own conclusion is the honest one, and it is why this is worth reading rather than cheering. No agent played beyond beginner level. None could put together a real army composition, none could defend a coordinated attack, none executed anything a human would call a strategy. Codex went 18-0 against the built-in AI and any experienced player would take it apart in ten minutes.
So hold both facts at once. A general-purpose coding agent, with no StarCraft training, beat a game's built-in bot eighteen times in a row at maximum difficulty, which nothing could do three years ago. And it did so while playing badly, at ten dollars a game, in a domain where the skill ceiling is six orders of magnitude above where it landed. Real-time strategy is the cleanest test of the thing agents are worst at, which is acting under a clock, and now we have a price tag for how badly they do it.
Comments