Anthropic Is Quietly Testing Whether You Notice Less Effort
"High" now means 10 out of 100 — the exact number "low" used to be. That's the finding a developer going by argofowl posted after digging into Claude Code's server-side configuration, and it put 127 points and an unusually bitter comment thread on Hacker News.
The specifics: Fable 5 sessions on Claude Code 2.1.236 and later are being enrolled, server-side, into an experiment that shrinks the reasoning-effort scale. Older client versions and Opus 5 sessions are untouched. Nothing in the changelog, and because it's an A/B test, only a slice of users are in it — which means two people on identical plans and identical settings are getting different models' worth of thinking, and neither of them knows.
This is not an isolated event. Anthropic runs live experiments on paying users through a remote feature-flag pipeline; previously documented tests include a Plan Mode variant that hard-capped output length and one that pulled Claude Code from a slice of Pro accounts entirely. And it lands on a pile we've been tracking for weeks: the Opus 5 feels-worse reports, the session token economics teardown, the "models are getting dumber on purpose" thesis. Those were vibes and inference. This one has a number.
To be clear about where the line is: every serving stack A/B tests, and effort routing per se is reasonable cost engineering. The problem is silently redefining a knob the user explicitly set. "High" is a promise with a meaning, and when the meter lies, developers respond the way they always do — by routing around you. The multi-vendor router ecosystem we keep covering is exactly where that trust goes when it leaves.
HN thread: https://news.ycombinator.com/item?id=49401549 original: https://x.com/argofowl/status/2091150597374537729
← Back to all articles
The specifics: Fable 5 sessions on Claude Code 2.1.236 and later are being enrolled, server-side, into an experiment that shrinks the reasoning-effort scale. Older client versions and Opus 5 sessions are untouched. Nothing in the changelog, and because it's an A/B test, only a slice of users are in it — which means two people on identical plans and identical settings are getting different models' worth of thinking, and neither of them knows.
This is not an isolated event. Anthropic runs live experiments on paying users through a remote feature-flag pipeline; previously documented tests include a Plan Mode variant that hard-capped output length and one that pulled Claude Code from a slice of Pro accounts entirely. And it lands on a pile we've been tracking for weeks: the Opus 5 feels-worse reports, the session token economics teardown, the "models are getting dumber on purpose" thesis. Those were vibes and inference. This one has a number.
To be clear about where the line is: every serving stack A/B tests, and effort routing per se is reasonable cost engineering. The problem is silently redefining a knob the user explicitly set. "High" is a promise with a meaning, and when the meter lies, developers respond the way they always do — by routing around you. The multi-vendor router ecosystem we keep covering is exactly where that trust goes when it leaves.
HN thread: https://news.ycombinator.com/item?id=49401549 original: https://x.com/argofowl/status/2091150597374537729
Comments