August 15, 2026AgentsCodingBenchmark

620 comments arguing that Opus 5 is smarter and worse to work with

The third-biggest story on Hacker News today is a blog post titled "Why does Opus 5 feel worse to work with?" — 671 points and 620 comments in eleven hours. The author is upfront that this is anecdote, calls their own explanation baseless speculation, and offers no data. Cover it anyway, because when six hundred engineers show up to agree about a feeling, the feeling is a datapoint about the industry even when it is not a datapoint about the model.

The complaint is specific and that is why it landed. Opus 5 benchmarks higher than Opus 4.7, 4.8 and Fable, and in daily use it makes assumptions and proceeds where the older models would have stopped and asked what you meant. The author's word is babysitting. Ambiguity used to produce a question; now it produces a confident branch, and if the branch is wrong you find out forty minutes later.

The proposed mechanism is the interesting claim. Benchmarks have a guaranteed right answer, so a model that makes bold usually-correct assumptions under ambiguity outscores one that stops to ask. Asking a clarifying question on a benchmark is a failed attempt. Asking one in a real repository is the correct move. If you optimize hard against the first environment you will systematically train out the behavior that matters in the second, and the more agentic and long-horizon the evals get, the stronger that gradient pulls.

Whether or not that is why Opus 5 behaves this way, the structural point is real and the timing is loud. Google's Gemini 3.7 Flash launch materials yesterday specifically advertised that it asks for clarification when intent is ambiguous — a vendor treating "knows when to stop and ask" as a headline feature. Meanwhile DarwinX and AutoDesign both published this week showing that a harness can be searched to fit a task. If the model's disposition under ambiguity is the problem, the harness is where you fix it, and nobody currently benchmarks disposition at all.

The gap this exposes is the useful takeaway. There is no widely used eval for "asked the right question instead of guessing," which means there is no pressure on anyone to be good at it. Post is at mun-logadan.github.io/why-does-opus-5-feel-worse.
← Previous
Mole caps your research agent's spend before it runs, and overshot by zero
Next →
Super User Daily: August 15, 2026
← Back to all articles

Comments

Loading...
>_