acceptodds
Under review as a conference paper at ICLR 2027

Locating Hidden Failures Makes Long-Horizon Agents More Reliable

Abstract

AI agents now carry out long tasks that take tens to hundreds of steps, yet they are judged almost entirely by whether their final answer passes. An outcome-based pass or fail signal cannot show where a trajectory went wrong or what harm the agent caused along the way. Identifying errors in an agent's thoughts and actions, step by step, would support robust monitoring, debugging, and dense, step-level rewards for training agents. This is challenging because humans cannot check long trajectories at scale, and the most visible errors are usually consequences of an earlier, hidden mistake. Existing process verification benchmarks cover only single-turn reasoning or short agentic tasks, often with artificially injected errors, and do not reflect how long-horizon agents fail in real settings, where they go wrong early, fake success, or cause real harm. We introduce Traverse, a benchmark for long-horizon agentic process verification consisting of real agent trajectories from software engineering, computer use, and bioinformatics data analysis, with human labels marking where the agent goes wrong. Current frontier models struggle to identify these process errors. Alongside the benchmark, we build a taxonomy of failure types, grouped into families, for long-horizon agents. We find that the mistakes that start a failure are most often reasoning and planning errors, and that agents take harmful, irreversible actions along the way, even in trajectories that pass the outcome check. We then train Scout, a B verifier, with supervised fine-tuning and reinforcement learning. It localizes mistakes more accurately than frontier models and transfers to a domain it never sees in training. At test time, it improves task success by percentage points on Terminal-Bench 2.0 and points on SWE-bench Verified. Our benchmark, failure taxonomy, and verifier training recipe offer a practical way to detect and localize failures in long-horizon agents and oversee them at scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.