From Eval Rollout To Process Supervision: Measuring And Curating LLM Agent Trajectories At Scale
Abstract
Coarse pass/fail evaluation can rank long-horizon LLM agents, but it provides little insight into how they solve tasks or why they fail. Two agents with the same outcome may exhibit fundamentally different processes: one failure may result from a minor omission near task completion, whereas another may stem from unfocused exploration, incorrect localization, or ineffective recovery. Likewise, successful trajectories are not equally valuable for training. A clean and well-grounded solution should not be treated as equivalent to one that succeeds through blind retries, editing before inspection, evaluator misuse, mechanical brute force, or excessive resource consumption. This raises a central question: which process-level signals should be used to identify trajectories with the greatest training value? Addressing this question is challenging because long-horizon coding tasks rarely admit a unique correct trajectory, while their final outcomes are often only partially verifiable. We introduce FGA, a fine-grained framework for evaluating LLM-agent trajectories beyond outcome verification. FGA combines automatically observable behavioral signals—such as tool-use sequences, read–edit order, retry dynamics, verification behavior, and resource consumption—with context-sensitive semantic process patterns, including grounded exploration, effective recovery, redundant churn, false verification, and premature convergence. The resulting evaluator provides interpretable assessments at both the step and trajectory levels, distinguishing clean success from wasteful or unreliable success, and productive failure from uninformative failure. We further study whether these process assessments can improve trajectory selection and weighting for supervised fine-tuning and reinforcement learning relative to outcome-only, rule-based, and generic LLM-judge baselines. By shifting evaluation from final correctness alone to evidence-grounded process quality, FGA aims to support agents that are not only more successful, but also more efficient, rigorous, and robust.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.