Recursive LLM-as-a-Judge for Long-Horizon Web-Agent Trajectories
Abstract
Recent advances in large language models (LLMs) have enabled web agents to perform increasingly complex, long-horizon tasks. Reliable evaluation of long-horizon web-agent trajectories requires effective evidence acquisition and joint reasoning over collected evidence. However, evidence acquisition is challenging when trajectories exceed the judge's context window, as exhaustive inspection is inefficient when evidence is sparse and fixed selection cannot adapt to uncertainties uncovered during evaluation. Moreover, evidence aggregation must then account for both cross-window dependencies and each rubric's logical structure. To address these challenges, we introduce RecursiveWebJudge, an LLM-as-a-judge framework inspired by recursive reasoning that couples adaptive evidence acquisition with rubric-specific aggregation. The framework first partitions trajectories into context-compatible windows and decomposes rubrics into verifiable requirements. For each selected window, the judge records local assessments and factual notes for these requirements in a shared evidence pool, using accumulated evidence and unresolved requirements to guide subsequent window selection and stopping. It then reasons over the accumulated evidence across windows to derive a final verdict, applying AND when all requirements must hold and ANY when a single qualifying instance suffices. We evaluate RecursiveWebJudge on AgentRewardBench, a benchmark of expert-annotated multimodal web-agent trajectories, using 3B, 7B, and 11B judge backbones to assess task success and non-progressing repetition. It improves average macro-F1 across task success and repetition detection by up to 26.6% over the strongest baseline, with its 7B judge surpassing full-trajectory Claude Code evaluation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.