Agent Error Dataset: 50,000 Failures, and the Fix Usually Works on the First Try
A failed agent run is usually thrown away with a zero reward. The Agent Error Dataset (AED, HF 46 upvotes) argues that's the most wasteful habit in agent training, because a failure carries everything you need to learn from it: what the agent saw, what it chose, and how the environment reacted.
AED collects 50,228 error-diagnosis pairs from 9,961 source tasks, spanning 33 environments, 19 harness families and 23 policy models. The original traces and execution metadata are kept, so anyone can re-diagnose a failure without rerunning the whole rollout. A five-stage pipeline gathers natural failures, generates a diagnosis and a proposed correction for each, and checks them against the recorded evidence. Where the environment can be replayed, the correction is tested against simply retrying the original action from the same checkpoint.
The headline number: across 3,062 matched replay pairs, the first proposed correction raises verifier pass rates from 18.4% to 51.1%. In other words, pinpointing the one wrong decision and swapping it out is nearly three times as effective as hoping a retry goes better. Fine-tuning Qwen3-8B on full diagnoses from 1,656 tasks lifts its agreement with teacher labels on the exact wrong step from 47.2% to 63.6%, beating the strongest prompted reference at 54.7%.
The practical takeaway: "find the step that went wrong" is a trainable skill, and a small model can learn it. That makes a cheap failure-diagnosis model a realistic component in your own agent stack, sitting beside the actor and reading its traces. Project page at agenterrordata.apodex.com.
Link: arxiv.org/abs/2609.40111
← Back to all articles
AED collects 50,228 error-diagnosis pairs from 9,961 source tasks, spanning 33 environments, 19 harness families and 23 policy models. The original traces and execution metadata are kept, so anyone can re-diagnose a failure without rerunning the whole rollout. A five-stage pipeline gathers natural failures, generates a diagnosis and a proposed correction for each, and checks them against the recorded evidence. Where the environment can be replayed, the correction is tested against simply retrying the original action from the same checkpoint.
The headline number: across 3,062 matched replay pairs, the first proposed correction raises verifier pass rates from 18.4% to 51.1%. In other words, pinpointing the one wrong decision and swapping it out is nearly three times as effective as hoping a retry goes better. Fine-tuning Qwen3-8B on full diagnoses from 1,656 tasks lifts its agreement with teacher labels on the exact wrong step from 47.2% to 63.6%, beating the strongest prompted reference at 54.7%.
The practical takeaway: "find the step that went wrong" is a trainable skill, and a small model can learn it. That makes a cheap failure-diagnosis model a realistic component in your own agent stack, sitting beside the actor and reading its traces. Project page at agenterrordata.apodex.com.
Link: arxiv.org/abs/2609.40111
Comments