Enabling Token-wise Importance Sampling in Asynchronous RLVR with Proximal KL Regularization
Abstract
In asynchronous reinforcement learning with verifiable rewards (RLVR), rollout generation overlaps with policy optimization, so training responses may come from earlier policy versions. Policy-gradient methods commonly reuse these responses with token-wise importance weighting, avoiding products of likelihood ratios across long responses. Under appropriate conditions, we prove that full future KL credit preserves the exact proximal optimizer as the token-wise update's unique stationary policy. This holds despite gradient bias, even when responses come from a mixture of older policies. The update requires neither full-response importance weights nor a fitted critic. We further establish convergence of tabular continuous-time population updates under persistent coverage and time-varying behavior policies. These results motivate **Proximal KL**, a practical asynchronous RLVR recipe that separates the proximal reference from the sampling policy and uses bounded future KL credit in place of ordinary PPO clipping. We evaluate **Proximal KL** on Qwen3-1.7B and Qwen3-4B, finding improvements over our DAPO control: higher off-policy MATH500 and MinervaMath accuracy at both tested learning rates for 1.7B and higher AIME24 and AIME25 mean@16 and pass@16 at 4B. Ablations show better preservation of AIME mean@16 as rollout lag increases and reduced MATH500 sensitivity to baseline perturbations. The horizon sweep favors a finite future window over both tested extremes, while fixed-policy measurements show higher correction variance at longer tested horizons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.