acceptodds
Under review as a conference paper at ICLR 2027

Bounding Token Divergence to Unlock Stable Reinforcement Learning for Large Language Models

Abstract

Decoupled reinforcement learning from verifiable rewards (RLVR) relies on separate inference and training engines that, despite sharing identical weights, assign different token-level probabilities due to floating-point non-determinism and stochastic expert routing. The resulting importance-sampling corruption destabilizes policy-gradient training, particularly for Mixture-of-Experts (MoE) models. We first introduce IcePop, a token-level ratio mask that eliminates extreme discrepancies between inference and training probabilities and restores training stability. Analysis reveals, however, that any constant ratio threshold over-masks low-probability tokens because the softmax partition function attenuates perturbations for frequent tokens while leaving rare tokens disproportionately vulnerable to them. We then propose KPop, which replaces the fixed ratio window with a symmetric binary KL divergence criterion. The binary KL naturally widens its tolerance at low token probability and tightens it at high probability, matching the heteroscedastic noise structure. The symmetric formulation further eliminates directional leakage inherent in one-sided constraints. Experiments on dense and MoE architectures across mathematical reasoning benchmarks show that KPop retains significantly more rare-token gradient signal than existing methods while improving both training stability and final evaluation accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.