acceptodds
Under review as a conference paper at ICLR 2027

Spurious Policy Shift: Do Not Be Too Conservative for Your Policy Optimization

Abstract

Reinforcement learning (RL) has become a key component of post-training for large language models (LLMs), where token-level clipping stabilizes off-policy optimization by controlling token-wise importance sampling (IS) ratios between the current policy and the rollout policy. However, we empirically identify a phenomenon termed spurious policy shift (SPS): a large shift in the sampled token's probability may simply reflect probability mass shifting among contextually interchangeable alternatives at the same decoding state, rather than a meaningful change in the model's generation policy. Furthermore, we observe that SPS-affected tokens occupy a substantial portion of upper-clipped positive advantage updates and suppress policy gain, suggesting that the standard token-level clipping is too conservative and may unnecessarily suppress positive policy updates. Building on this observation, we propose SPS-aware clipping for policy optimization (SCPO) that identifies SPS through probability mass shifting between the sampled token and its contextually interchangeable candidate tokens under the old and new policies, and selectively relaxes clipping for such tokens to avoid suppressing positive policy updates. Experimental results show that our method outperforms state-of-the-art methods on multiple math reasoning benchmarks without compromising training stability. Our code is available at: https://anonymous.4open.science/r/scpo

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.