Sonnet 5.5 Beats Opus 5.5 at Terminal Work, for a Fraction of the Price
The middle child just beat the flagship. Anthropic shipped Claude Sonnet 5.5 on Monday, September 28, and on Terminal-Bench 4.0 it scores 70.6%. Opus 5.5 scores 66.4%. Sonnet 5, the model it replaces, scored 10.3%. That last number is not a typo, and it is the whole story: a seven-fold jump on long terminal sessions in one generation.
The pricing is $2 per million input tokens and $10 per million output, with cache reads at $0.20. Anthropic says it runs 30%+ faster than Sonnet 5 and costs up to 30% less per task, which matters more than the per-token price because agents are billed by the loop, not the call. OSWorld 2.1 lands at 80.1% against Opus 5.5's 81.8%, GDPval-AA is basically a tie at 1844 vs 1846, and on CursorBench and FrontierCode Opus still leads by a few points. It is also the first Sonnet to beat PokΓ©mon Red from screenshots, which sounds like a gimmick until you remember it is a hundreds-of-hours computer-use task with no API.
Two details worth noticing. First, Anthropic published a cost-per-task figure, not just a price sheet. That is the number every agent builder actually needs, and it is still rare. Second, this is the first Sonnet under the same cyber safeguards as Opus and Fable, plus classifiers against reasoning extraction for distillation. A mid-tier model that can do this much in a terminal is now treated as a dual-use tool.
The practical read: for most agent loops, the default just changed. Opus stays the pick for the hardest coding and design work, but anyone running long, well-scoped tasks at volume should rerun their evals on claude-sonnet-5-5 this week. A new Haiku is promised in the coming weeks, which would reshuffle the router tier again.
Link: anthropic.com/claude-sonnet-5-5
← Back to all articles
The pricing is $2 per million input tokens and $10 per million output, with cache reads at $0.20. Anthropic says it runs 30%+ faster than Sonnet 5 and costs up to 30% less per task, which matters more than the per-token price because agents are billed by the loop, not the call. OSWorld 2.1 lands at 80.1% against Opus 5.5's 81.8%, GDPval-AA is basically a tie at 1844 vs 1846, and on CursorBench and FrontierCode Opus still leads by a few points. It is also the first Sonnet to beat PokΓ©mon Red from screenshots, which sounds like a gimmick until you remember it is a hundreds-of-hours computer-use task with no API.
Two details worth noticing. First, Anthropic published a cost-per-task figure, not just a price sheet. That is the number every agent builder actually needs, and it is still rare. Second, this is the first Sonnet under the same cyber safeguards as Opus and Fable, plus classifiers against reasoning extraction for distillation. A mid-tier model that can do this much in a terminal is now treated as a dual-use tool.
The practical read: for most agent loops, the default just changed. Opus stays the pick for the hardest coding and design work, but anyone running long, well-scoped tasks at volume should rerun their evals on claude-sonnet-5-5 this week. A new Haiku is promised in the coming weeks, which would reshuffle the router tier again.
Link: anthropic.com/claude-sonnet-5-5
Comments