When Policies Drift, Prefixes Matter: Prefix-Normalized Policy Optimization for LLM Reinforcement Learning
Abstract
In LLM reinforcement learning, the behavior policy can lag behind the current policy through asynchronous generation and updates or repeated rollout reuse; numerical differences between training and inference add further mismatch. Prox- imal methods such as PPO use local likelihood ratios to correct for conditional action-probability mismatch, but not explicitly for differences in prefix visitation induced by earlier decisions. In autoregressive generation, each prefix has a unique path, yielding an exact state–action change-of-measure ratio at each position: the product of local likelihood ratios through that position. However, this ratio can grow or decay rapidly with prefix length, hindering long-sequence optimization. We propose Prefix-Normalized Policy Optimization (PNPO), which geometrically averages local likelihood ratios over each causal prefix and constrains updates with position-adaptive clipping. Under equal optimizer-update budgets, PNPO outperforms GRPO and GSPO in mean Pass@1 across six mathematical reason- ing benchmarks across different policy-lag settings, with a larger advantage under high policy lag. With four PPO epochs, PNPO achieves 52.9% mean Pass@1, compared with 50.9% for GSPO and 50.7% for GRPO. PNPO and GSPO remain stable, whereas GRPO declines late in training. PNPO also matches its one-epoch performance using one quarter as many fresh rollout batches.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.