Agent Error Dataset: Scaling 50,000 Error–Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training
Abstract
LLM agents learn from demonstrations and reward feedback, yet an unsuccessful rollout does not specify which decision should change. How can failed expe- rience become reusable supervision for diagnosing errors and improving agent behavior? We introduce the Agent Error Dataset (AED), comprising 50,228 error–diagnosis pairs across 33 environments, 19 harness families and 23 policy models in text-based agent systems for failure analysis and post-training. Our five-stage Agentic Error-to-Training (AET) pipeline generates trace-grounded diagnoses and proposed corrections, tests corrections against original-action re- tries from the same checkpoint where replay is supported, and constructs separate training views for diagnosis and recovery. Across 3,062 matched replay pairs, first-proposal corrections improve verifier pass rates by 32.7 percentage points over their original-action controls. Training on up to 1,656 source tasks from a separately frozen diagnosis release raises Qwen3-8B’s exact-step agreement with internal teacher labels from 47.2% to 63.6%, averaged over three seeds. In a single-seed actor recipe comparison, action-only repair training scores 6.67 per- centage points higher on WebShop-lite than success-only training, with gains that depend on the environment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.