acceptodds
Under review as a conference paper at ICLR 2027

Process-Supervised Learning for Agentic Search via Dual-Granularity Advantage

Abstract

Reinforcement learning (RL) has emerged as a promising paradigm for optimizing search agents in complex reasoning tasks. However, traditional outcome-based RL often relies on sparse terminal rewards and suffers from credit assignment challenges, resulting in gradient starvation and high-variance policy updates. This limits optimization stability and can lead to process hallucinations, where agents achieve correct answers via flawed reasoning or redundant retrieval. While recent process-aware approaches attempt to mitigate these issues, they often lack on-policy exploration capability or collapse step-level credit into coarse trajectory-level optimization. To address these limitations, we propose ProSearch, a dual-granularity process-supervised RL framework. Leveraging a Monte Carlo Tree Search (MCTS)-derived Process Reward Model (PRM), ProSearch introduces a novel dual-granularity advantage mechanism that combines dense step-level signals with sparse outcome rewards. This provides precise credit assignment for specific actions, mitigating gradient starvation while theoretically improving the Signal-to-Noise Ratio (SNR). Extensive experiments on five multi-hop QA benchmarks demonstrate that ProSearch achieves overall superior performance compared to strong baselines, particularly on complex long-horizon tasks. Further fine-grained analysis reveals that ProSearch not only stabilizes training dynamics but also significantly reduces logical fallacies and redundant exploration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.