UndoBench: Agents Finish 84% of Tasks and Recover From 47% of Faults
Every enterprise agent benchmark measures whether the task got done. Almost none measure what happens when a tool call fails halfway through. UndoBench, posted Monday, separates the two, and the gap is the finding. Across 5,760 executions on held-out workflows, nominal competence was 83.54%. Conditional recovery success, meaning the agent got the task right after a fault was injected, was 46.72%. The same agents that look competent lose nearly half their reliability the moment something goes wrong.
The benchmark has 36 base workflows and 36 fault scenarios across 8 enterprise domains, with counterfactual paired trials run under identical seeds so each fault run has a clean twin. Wire-level effect history and environment-state oracles record what the agent actually did to the outside world, not what it reported. That is how the paper gets its ugliest number: naive retry produced duplicate external effects in 53.33% of trials. A payment sent twice, a ticket filed twice, an email sent twice. Two open-weight models, two frameworks and three recovery paradigms were tested, and commercial API models reproduced the same competence-recovery split.
The useful part is that recovery depends on where in the operation the fault lands. Before any mutation, every method does about the same and nothing duplicates. During partial mutation, naive retry, per-call idempotency and zero-privilege journaling all collapse on composite workflows. After the write committed but before the acknowledgment arrived, verification and server-side idempotency substantially improve safety. So the fix is not in the prompt. It is in whether the tool on the other end can tell the agent "you already did that".
This lines up with the week's enforcement thread. The agent's intent is fine. The receipt is missing. UndoBench gives that gap a number, and 53% duplicate side effects is the number to put in front of anyone wiring an agent to a payments API.
Link: arxiv.org/abs/2610.05622
← Back to all articles
The benchmark has 36 base workflows and 36 fault scenarios across 8 enterprise domains, with counterfactual paired trials run under identical seeds so each fault run has a clean twin. Wire-level effect history and environment-state oracles record what the agent actually did to the outside world, not what it reported. That is how the paper gets its ugliest number: naive retry produced duplicate external effects in 53.33% of trials. A payment sent twice, a ticket filed twice, an email sent twice. Two open-weight models, two frameworks and three recovery paradigms were tested, and commercial API models reproduced the same competence-recovery split.
The useful part is that recovery depends on where in the operation the fault lands. Before any mutation, every method does about the same and nothing duplicates. During partial mutation, naive retry, per-call idempotency and zero-privilege journaling all collapse on composite workflows. After the write committed but before the acknowledgment arrived, verification and server-side idempotency substantially improve safety. So the fix is not in the prompt. It is in whether the tool on the other end can tell the agent "you already did that".
This lines up with the week's enforcement thread. The agent's intent is fine. The receipt is missing. UndoBench gives that gap a number, and 53% duplicate side effects is the number to put in front of anyone wiring an agent to a payments API.
Link: arxiv.org/abs/2610.05622
Comments