acceptodds
Under review as a conference paper at ICLR 2027

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Abstract

Reinforcement learning is widely used to improve reasoning in large language models. A central design choice is how much policy change to allow at each token of a response. Uniform token-level thresholds ignore both token position and preceding policy divergence. Yet early changes can affect longer continuations, and divergence can accumulate along the prefix. To address this, we introduce CPPO (Cumulative Prefix Divergence Policy Optimization), a token-level masking rule that combines position-weighted divergence with a cumulative prefix budget. Divergence at earlier positions receives greater weight, and the prefix budget limits the cumulative weighted divergence. We derive a finite-horizon policy improvement bound under exact token divergence constraints. With a sufficiently small prefix budget, it is tighter than the bound for uniform token-level limits at the same nominal threshold. Experiments on mathematical reasoning show that CPPO outperforms the evaluated baselines across model sizes. Ablation studies support the effectiveness of both components.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.