acceptodds
Under review as a conference paper at ICLR 2027

Tracepidia: Diagnosing Failures Behind Failures in Agent Trajectories

Abstract

Agents can fail long-horizon tasks not only because of the model's own limitations but also because of the tools, the execution environment, or other components of the system. Correcting such a failure requires identifying the responsible component and how its behavior affected the outcome. Existing approaches label the responsible agent and the error step from the agent's trajectory. Such a label cannot tell a tool fault from a model fault, and the trajectory does not show when a tool or the context manager changes what the model receives. To this end, we introduce Tracepidia, which consists of a benchmark, an auditor, and a toolkit for agent failure diagnosis. The benchmark contains 2,332 samples on 71 Terminal-Bench 2.1 tasks. Each sample injects one fault into a passing trajectory in which every step succeeds, so its responsible component, failure category, and error step are known. The auditor names these three from the trajectory and the telemetry that the agent harness emits while it runs, and the toolkit records this telemetry without changes to the agent. On held-out tasks, the auditor raises macro-F1 over failure categories from 0.315 to 0.424 against a direct judge that uses the same LLM. The gain is largest on samples that still pass the verifier, where no failed test shows what went wrong. Across 890 sessions of ten agent configurations, the auditor finds a failure in 57.10% of the sessions that pass the verifier. A stronger model lowers the share of sessions with a failure, and higher reasoning effort does not.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.