LongRCA Bench: Root-Cause Localization in Long-Horizon Failed Agent Trajectories
Abstract
Failure attribution is a critical component of the failure-driven agent improve- ment loop: identifying which workflow role introduced a decisive error and when can turn failed executions into actionable feedback for repair, retry, and workflow refinement. However, existing benchmarks are dominated by short execution his- tories. In Who&When, 71% of trajectories contain at most 10 recorded steps, a regime in which the visible erroneous step is often close to the final failure and failure attribution can collapse toward local failure detection. Long-horizon failures create a different diagnostic problem: an early error may be propagated, reused, validated, or left uncorrected across dozens or hundreds of subsequent steps before the execution terminates. We introduce LONGRCA BENCH, com- prising 1,140 complete naturally failed trajectories from five sources, all human- annotated with responsible roles and earliest decisive root-cause steps. Under the same DeepSeek-V4-Flash backbone and scoring protocol, five existing meth- ods achieve 30.4–36.4% exact root-step accuracy on Who&When but only 2.8– 13.2% on LONGRCA BENCH, exposing a substantial gap between detecting nearby errors and tracing distant error origins. Motivated by the benchmark’s challenges, we propose Root-Cause Trajectory Attribution (RCTA), a training- free method that combines hierarchical trajectory summarization, original-record retrieval, and upstream instruction–execution comparison to distinguish inherited errors from newly introduced ones. RCTA achieves 41.8% exact root-step accu- racy on Who&When and 24.1% on LONGRCA BENCH, compared with 36.4% and 13.2% for the strongest evaluated baselines, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.