Learning at the Branch Point: State-Local Policy Optimization for Search Agents
Abstract
Reinforcement learning (RL) has emerged as an effective approach for training large language model (LLM)-based search agents on knowledge-intensive tasks requiring iterative information seeking. However, existing RL methods typically optimize at the trajectory level, where credit assignment based primarily on whole-trajectory outcomes can obscure the individual contributions of useful and redundant search actions. To address this limitation, we propose State-Local Policy Optimization (SLPO), which shifts the optimization unit from whole trajectories to individual actions under the same search history. At each candidate state, SLPO executes and scores alternative actions, deriving local advantages from their relative rewards for action-level updates. To explore diverse search directions while limiting rollout cost, we structure exploration as a semantic tree, clustering similar actions and expanding only cluster representatives while retaining all sampled alternatives for local comparison. We use precomputed evidence-to-document mappings to reward actions that retrieve new evidence without an LLM judge during training. We further prioritize states with greater within-state reward dispersion for more informative local comparisons. Across seven QA benchmarks, SLPO improves answer accuracy (COVER-EM) over the strongest baseline by 3.54 and 3.84 percentage points for 4B- and 8B-parameter models, respectively, while using comparatively few searches. These results demonstrate the effectiveness of state-local credit assignment for training LLM-based search agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.