DeltaPRM: Rollout-Free Automatic Process Rewards via Answer Likelihood Gain for Retrieval Agents
Abstract
Retrieval-augmented language agents have become a central paradigm for knowledge-intensive question answering, yet they are typically trained with outcome rewards alone, leaving individual search interactions unsupervised. Process reward models (PRMs) address this gap, but their step-level labels are expensive to obtain: human annotation and LLM-as-a-judge labels both demand substantial resources, while Monte-Carlo (MC) estimation, the dominant annotation-free alternative, becomes prohibitive for retrieval agents, since each per-step continuation triggers additional interaction with the search environment. We introduce DeltaPRM, a process reward that requires neither extra rollouts nor extra environment interaction. DeltaPRM scores each retrieval call and the reasoning that follows it by how much the information it brings raises the model's likelihood of the correct final answer — the delta between steps. Labeling thus reduces to a single forward pass per step and relies only on outcome supervision already available in any QA dataset. We train a PRM to predict the discounted return of this label signal from the reasoning prefix alone, no longer requiring access to the ground truth final answer. Across multi-hop agentic search benchmarks, we test DeltaPRM on best-of- answer selection, step-level guided generation (greedy decoding, beam search, and MC tree search), step-level verification on AgentProcessBench, and online RL, where it consistently matches or surpasses substantially more expensive baselines, including LLM-as-a-judge PRMs. At the same time, DeltaPRM reduces labeling cost by and relative to MC estimation and LLM-as-a-judge, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.