Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
Abstract
Reinforcement learning is widely used to improve reasoning in large language models. A central design choice is how much policy change to allow at each token of a response. Uniform token-level thresholds ignore both token position and preceding policy divergence. Yet early changes can affect longer continuations, and divergence can accumulate along the prefix. To address this, we introduce CPPO (Cumulative Prefix Divergence Policy Optimization), a token-level masking rule that combines position-weighted divergence with a cumulative prefix budget. Divergence at earlier positions receives greater weight, and the prefix budget limits the cumulative weighted divergence. We derive a finite-horizon policy improvement bound under exact token divergence constraints. With a sufficiently small prefix budget, it is tighter than the bound for uniform token-level limits at the same nominal threshold. Experiments on mathematical reasoning show that CPPO outperforms the evaluated baselines across model sizes. Ablation studies support the effectiveness of both components.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.