Two Papers, One Finding: Agents Collect the Evidence and Then Ignore It
Two papers on this week's HuggingFace board land the same punch from different sides. The first, From Evidence to Action: How Tool-Using Agents Fail (arXiv 2610.07753, 34 upvotes), introduces SafeActBench, 656 cases across six operational domains and five protocols that move from "judge this action" to "investigate and decide not to act" to single-action and multi-action workflows. A provenance-bound Evidence Ledger records what the agent knew and when it acted. Across ten model-harness configurations, the authors find that strong static judgment coexists with much weaker interactive execution, and that failures start before execution: agents stop investigating early, or act before the required evidence exists. Once the evidence is in hand, single actions are usually fine. Multi-step workflows then break on unresolved prerequisites.
The second paper, Judged Useless, Queried Anyway (arXiv 2610.06191), is narrower and sharper. In a retrieval environment with a deliberately failing source, seven agents label that source's results useless 97 to 100 percent of the time when asked. Most of them keep querying it anyway. Prompt cues change when they stop but not what they stop on. Giving permission to answer from memory or turning on a reasoning mode produces early stops regardless of evidence. A stated budget pushes the 7B to 8B models to run until the deadline. Stopping rules and call costs in the prompt are followed at most partly. The only thing that worked: a harness-enforced integration step that forces the agent to answer after five consecutive results it judged useless. That raised failing-source success for every model, held the stopping point fixed when the budget doubled, and survived a pre-registered replication on 300 fresh questions.
Read together, the claim is that the bottleneck in tool-using agents is not perception and not knowledge. It is the link between a judgment the model has already made and the action it takes next, and that link does not hold inside the model. It holds when the harness closes it. This is the enforcement-over-instruction thread with two more controlled measurements, and the second paper's pre-registration is the first time one of these results has come with a replication built in.
Links: arxiv.org/abs/2610.07753, arxiv.org/abs/2610.06191
← Back to all articles
The second paper, Judged Useless, Queried Anyway (arXiv 2610.06191), is narrower and sharper. In a retrieval environment with a deliberately failing source, seven agents label that source's results useless 97 to 100 percent of the time when asked. Most of them keep querying it anyway. Prompt cues change when they stop but not what they stop on. Giving permission to answer from memory or turning on a reasoning mode produces early stops regardless of evidence. A stated budget pushes the 7B to 8B models to run until the deadline. Stopping rules and call costs in the prompt are followed at most partly. The only thing that worked: a harness-enforced integration step that forces the agent to answer after five consecutive results it judged useless. That raised failing-source success for every model, held the stopping point fixed when the budget doubled, and survived a pre-registered replication on 300 fresh questions.
Read together, the claim is that the bottleneck in tool-using agents is not perception and not knowledge. It is the link between a judgment the model has already made and the action it takes next, and that link does not hold inside the model. It holds when the harness closes it. This is the enforcement-over-instruction thread with two more controlled measurements, and the second paper's pre-registration is the first time one of these results has come with a replication built in.
Links: arxiv.org/abs/2610.07753, arxiv.org/abs/2610.06191
Comments