Fable 5.1 Lands: Cheaper Loops, Longer Runs, Looser Filters
Anthropic shipped Claude Fable 5.1 and Mythos 5.1 on September 1. Same model, two trims: Fable is generally available with standard safeguards, Mythos goes only to vetted cybersecurity and life-science organizations with the filters dialed down. That two-tier structure debuted with Fable 5; now it's clearly the permanent shape of Anthropic's releases.
The number that matters most for agent builders isn't a benchmark, it's the invoice. Cache reads drop 75% to $0.25 per million tokens, which Anthropic says works out to roughly 25% cheaper for typical workloads and up to 45% cheaper for highly agentic ones. Agent loops re-read the same context hundreds of times, so cache pricing IS agent pricing. Headline rates stay at $10 in, $50 out.
The benchmarks tell one consistent story: long-horizon work. Terminal-Bench-Science more than doubles from 24.7% to 52.6%, agentic coding on Terminal-Bench 4.0 goes from 42% to as high as 60.9%, and the flagship anecdote is a 38-hour unattended run that diagnosed a published result as a label artifact. The pitch has quietly shifted from "smart in a chat window" to "trustworthy alone overnight."
The safeguards changes are almost a product launch of their own: cyber false positives down 60%, vulnerability discovery now allowed (exploit writing still isn't), biology refusals down 85% for benign queries. After a season where every lab tightened filters, Anthropic is betting that precision, not strictness, is the differentiator. We covered the Mythos gating logic when it debuted: https://clauday.com/article/d28ddc6f-9894-46c3-92e8-b858bb43793d
Announcement: https://www.anthropic.com/claude-fable-and-mythos-5-1
← Back to all articles
The number that matters most for agent builders isn't a benchmark, it's the invoice. Cache reads drop 75% to $0.25 per million tokens, which Anthropic says works out to roughly 25% cheaper for typical workloads and up to 45% cheaper for highly agentic ones. Agent loops re-read the same context hundreds of times, so cache pricing IS agent pricing. Headline rates stay at $10 in, $50 out.
The benchmarks tell one consistent story: long-horizon work. Terminal-Bench-Science more than doubles from 24.7% to 52.6%, agentic coding on Terminal-Bench 4.0 goes from 42% to as high as 60.9%, and the flagship anecdote is a 38-hour unattended run that diagnosed a published result as a label artifact. The pitch has quietly shifted from "smart in a chat window" to "trustworthy alone overnight."
The safeguards changes are almost a product launch of their own: cyber false positives down 60%, vulnerability discovery now allowed (exploit writing still isn't), biology refusals down 85% for benign queries. After a season where every lab tightened filters, Anthropic is betting that precision, not strictness, is the differentiator. We covered the Mythos gating logic when it debuted: https://clauday.com/article/d28ddc6f-9894-46c3-92e8-b858bb43793d
Announcement: https://www.anthropic.com/claude-fable-and-mythos-5-1
Comments