Just Reuse the Rollout Log-Probs: Stable Mixture-of-Experts Reinforcement Learning with a Vocabulary-KL Clip
Abstract
In reinforcement learning of large language models, responses are typically generated by a fast inference engine while gradients are computed by a different training backend, and the two disagree on the same tokens because of numerical and, for mixture-of-experts models, expert-routing differences. The PPO/GRPO ratio clip that stabilizes these updates acts on the sampled token's probability ratio: a one-dimensional proxy that neither measures how far the full next-token distribution moved nor, under trainer recompute, reflects which policy generated the data. We find that a double clip stabilizes these updates, and that with it the cheapest correction suffices. VoKL-Clip clips each token by a top- bucket KL between the rollout and current policies (computed for free from the top- log-probs the inference engine returns on request), removing drifted tokens from the gradient by stop-gradient, and keeps a wide ratio clip alongside it. The bucket coarsening yields the sharpest full-vocabulary total-variation bound available from top- information, yet the vocabulary-KL clip provably cannot control importance-weight moments, which the ratio clip bounds; dropping either clip costs 7–11 AIME24 points. With both, VoKL-Clip can just reuse the rollout's own log-probs as the behavior policy: the importance ratio is faithful to the sampler, with no recompute and no routing replay. On Qwen3-30B-A3B trained on DAPO-Math-17k, VoKL-Clip reaches 64.5% AIME24 (ave@32)—versus 58.1% for DAPO and 57.2% for GSPO, each at its overall best rollout↔training mismatch-correction level—leads across the reasoning-benchmark suite, and on AIME24 exceeds every baseline at every correction level at lower correction cost, with stable training dynamics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.