Opus 5.5 is cheaper than Opus 5 and beats it everywhere
Anthropic shipped Claude Opus 5.5 on September 22, and the interesting part is not the benchmark table. It is the price. Input dropped from $5 to $4 per million tokens, output from $25 to $20, and cache reads from $0.50 to $0.20. That last one is a 60% cut and it matters more than the other two combined, because a long-running agent spends most of its tokens re-reading context it already paid to write. Anthropic's own estimate is about 40% cheaper on typical workloads, with 30% faster output.
The scores went up at the same time. Terminal-Bench 4.0 at 66.4%, FrontierCode v1.1 at 54.4%, CursorBench 4.0 at 57.8%, and OSWorld 2.0 at 81.8% for computer use. On GDPval-AA v2.1, the knowledge-work eval, it hits 1846 Elo against Fable 5.1's 1735 and Opus 5's 1708. One early tester reportedly pushed a 680,000-line code migration through in under a day. Take that anecdote for what it is, but the shape of it is right: the jobs people are handing these models are getting measured in repositories now, not files.
The model ID is claude-opus-5-5 and it is live on the Claude Platform, AWS, Google Cloud and Azure. Effort levels run low through max, plus a fast mode that buys 2.5x speed at a higher price. The anti-distillation safeguard on preserved thinking is still there.
Buried in the safety section is the number a lot of people building agent infrastructure should read twice: Opus 5.5 attempts sandbox escapes 85% less often than Opus 5, measured across an automated behavioral audit of 2,000-plus scenarios. That is a vendor publishing a rate for a failure mode that most vendors still pretend does not have a rate. Whatever you think of the absolute number, having one at all changes the conversation with your security team.
https://www.anthropic.com/claude-opus-5-5
← Back to all articles
The scores went up at the same time. Terminal-Bench 4.0 at 66.4%, FrontierCode v1.1 at 54.4%, CursorBench 4.0 at 57.8%, and OSWorld 2.0 at 81.8% for computer use. On GDPval-AA v2.1, the knowledge-work eval, it hits 1846 Elo against Fable 5.1's 1735 and Opus 5's 1708. One early tester reportedly pushed a 680,000-line code migration through in under a day. Take that anecdote for what it is, but the shape of it is right: the jobs people are handing these models are getting measured in repositories now, not files.
The model ID is claude-opus-5-5 and it is live on the Claude Platform, AWS, Google Cloud and Azure. Effort levels run low through max, plus a fast mode that buys 2.5x speed at a higher price. The anti-distillation safeguard on preserved thinking is still there.
Buried in the safety section is the number a lot of people building agent infrastructure should read twice: Opus 5.5 attempts sandbox escapes 85% less often than Opus 5, measured across an automated behavioral audit of 2,000-plus scenarios. That is a vendor publishing a rate for a failure mode that most vendors still pretend does not have a rate. Whatever you think of the absolute number, having one at all changes the conversation with your security team.
https://www.anthropic.com/claude-opus-5-5
Comments