SP-GRPO: Shared-Prefix Group Relative Policy Optimization for Multi-Turn Agents
Abstract
Reinforcement learning post-training, notably the critic-free Group Relative Policy Optimization (GRPO), has become a cornerstone for enhancing the reasoning capabilities of large language models. However, extending its group comparison paradigm to long-horizon, multi-turn agent tasks remains challenging, where sparse terminal rewards fail to provide fine-grained, turn-level credit assignment. Existing methods typically attempt isolated fixes by either modifying credit assignment through turn-level advantages calculated across independently rolled-out trajectories, or altering rollout structures via tree expansions. By treating rollout dynamics and credit assignment as decoupled components, these approaches suffer from a core structural misalignment, introducing severe context-confounding bias. To resolve this structural misalignment, we propose Shared-Prefix Group Relative Policy Optimization (SP-GRPO), which seamlessly aligns single-path environment continuation with turn-level counterfactual evaluation. By concentrating the sampling budget on multi-action branching from strictly identical historical states, SP-GRPO evaluates candidate actions under unbiased context, computes exact process advantages via instantaneous feedback, and advances the environment through a single active continuation path. Theoretically, we prove that SP-GRPO strictly preserves the on-policy trajectory distribution, eliminates context confounding bias, and achieves provable variance reduction over cross-trajectory estimators. Empirically, extensive experiments across seven single-hop and multi-hop QA benchmarks show that SP-GRPO significantly outperforms trajectory-level GRPO, tree-based approaches, and existing step-level baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.