TraceACE: Analysis, Clustering, and Explanation of Failures in LLM-Agent Traces
Abstract
Debugging an LLM agent from its final answer is underdetermined: the answer does not reveal which span introduced an error, whether failures across traces share one fix, or whether a fluent root-cause narrative matches recorded execution. This paper presents TRACEACE, a collection-level workflow that turns heterogeneous observability traces into localized issues, interpretable same-fix clusters, evidence-constrained root-cause hypotheses, and an engineering report. Canonicalization preserves source span identifiers; taxonomy-conditioned scanners run alongside optional rule-based anchors and a structural channel; transparent signatures and a conservative merge form same-fix clusters, each investigated via one representative; and every hypothesis quotation is checked against the workspace before replay on secondary traces, escalating on disagreement. Across TRAIL and AgentErrorBench with eight LLM backbones and seven baselines, TRACEACE attains the highest joint accuracy on TRAIL GAIA under every backbone, with a 24.2-point mean gain over the strongest baseline, and on TRAIL SWE under seven of eight; on critical root-cause attribution it matches the strongest per-trace attributors. Ablations show the largest losses when specialized scanners are replaced by a unified one (up to 3.3 joint points) or cross-trace validation is removed (about 4 points of root accuracy). TRACEACE uses less than half the per-trace tokens of AgentDebug, and cluster-level investigation raises root accuracy over per-trace investigation on both AgentErrorBench environments (16.2 vs. 14.9 and 20.4 vs. 18.4). These are single-run point estimates; same-fix gold labels and interventional validation remain unavailable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.