acceptodds
Under review as a conference paper at ICLR 2027

SR-Search: Structured Routing of Search and Refine Information Gains for Fine-Grained Credit Assignment

Abstract

Reinforcement learning (RL) has become a dominant paradigm for enhancing retrieval-augmented generation (RAG) in large language models. However, existing frameworks such as Search-R1 rely on trajectory-level outcome rewards assigned to all generated tokens along an execution path. This coarse and sparse signal cannot credit intermediate decisions, and in complex multi-hop question answering (QA) it fails to isolate whether a specific retrieval acquired relevant information or whether the subsequent refinement preserved and used that evidence. To address these limitations, we propose SR-Search, a fine-grained credit-assignment framework for search agents trained via Group Relative Policy Optimization (GRPO). Within each valid search–refine cycle, the policy evaluates the log-probability of the ground-truth answer before search, after retrieval, and after refinement; the resulting likelihood differences define Search Information Gain (SIG) and Refinement Information Gain (RIG), respectively, without any auxiliary reward model. SR-Search routes SIG to search actions and RIG to refinement actions, and further introduces a short-horizon delayed advantage that propagates refinement gains backwards to optimize the preceding retrieval decisions. Across seven QA benchmarks, SR-Search outperforms outcome-reward baseline AutoRefine by 3.7 points and the strongest step-reward baseline GiGPO by 2.1 points. These results demonstrate that temporally localized credit assignment can improve both retrieval and evidence refinement in RL-based search agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.