Spurious Policy Shift: Do Not Be Too Conservative for Your Policy Optimization
Abstract
Reinforcement learning (RL) has become a key component of post-training for large language models (LLMs), where token-level clipping stabilizes off-policy optimization by controlling token-wise importance sampling (IS) ratios between the current policy and the rollout policy. However, we empirically identify a phenomenon termed spurious policy shift (SPS): a large shift in the sampled token's probability may simply reflect probability mass shifting among contextually interchangeable alternatives at the same decoding state, rather than a meaningful change in the model's generation policy. Furthermore, we observe that SPS-affected tokens occupy a substantial portion of upper-clipped positive advantage updates and suppress policy gain, suggesting that the standard token-level clipping is too conservative and may unnecessarily suppress positive policy updates. Building on this observation, we propose SPS-aware clipping for policy optimization (SCPO) that identifies SPS through probability mass shifting between the sampled token and its contextually interchangeable candidate tokens under the old and new policies, and selectively relaxes clipping for such tokens to avoid suppressing positive policy updates. Experimental results show that our method outperforms state-of-the-art methods on multiple math reasoning benchmarks without compromising training stability. Our code is available at: https://anonymous.4open.science/r/scpo
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.