acceptodds
Under review as a conference paper at ICLR 2027

When Clipping Becomes a Learning Signal: A Permutation Baseline for Policy Optimization

Abstract

Reinforcement learning with verifiable rewards (RLVR) has improved mathematical reasoning in large language models. Methods such as Group Relative Policy Optimization (GRPO) use clipping to stabilize training by discouraging large policy changes. However, we identify a component of the clipped objective whose gradient acts as an additional learning signal outside the clipping interval and can counteract reward-guided updates. This component equals the average clipped objective over all within-group reward permutations, which we call the permutation baseline. In this work, we propose Permutation-Corrected Policy Optimization (PerCPO), which removes this source of interference with reward-guided updates by subtracting this baseline from the original clipped objective in closed form. For binary rewards, the resulting updates follow the direction indicated by rewards and are naturally attenuated outside the clipping interval, restoring gradients suppressed by clipping. We further show that the PerCPO objective increases along a path that improves accuracy with minimal Kullback–Leibler divergence from the sampling policy, even where the standard clipped objective plateaus. Extensive experiments on mathematical reasoning benchmarks demonstrate the effectiveness of PerCPO. For instance, on Qwen2.5-Math-7B, PerCPO outperforms GRPO by 3.47% in average accuracy across seven benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.