Judging the Path, Not Just the Answer: Turn-Level Verifiers for Deep Research Agents
Abstract
Autonomous deep research agents tackle complex, open-ended questions by sequentially reasoning and interacting with external environments across extended horizons. The prevailing reinforcement learning paradigm relies on terminal outcome rewards and encounters a fundamental turn-level credit assignment bottleneck. For failed trajectories, terminal penalties uniformly suppress all intermediate actions without distinguishing productive exploratory steps from genuine errors. On challenging queries where all rollouts fail, the learning signal vanishes. While process-level reward paradigms offer finer-grained feedback, they typically evaluate actions against ground-truth answers, limiting their ability to credit intermediate progress along unsuccessful paths and precluding their deployment during test-time search. To this end, we propose TPV (Turn-level Process Verifier), an unprivileged process verifier that evaluates each intermediate action purely from the interaction context. TPV is trained to predict turn-level grades assigned by an offline, privileged teacher judge that evaluates each action independently of final trajectory success. Empirical evaluations show that TPV reliably discriminates informative from redundant turns even inside failed rollouts, outperforming outcome-based baselines and transferring across diverse model scales. Consequently, our unified verifier serves a dual purpose: it delivers dense, fine-grained credit assignment during reinforcement learning, bypassing the zero-gradient dilemma on all-fail queries, and acts as an effective trajectory reranker at inference time. Extensive experiments on five challenging deep research benchmarks demonstrate that integrating TPV into RL training consistently boosts task success rates over competitive baselines, while deploying it as a test-time reranker yields substantial additional accuracy gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.