acceptodds
Under review as a conference paper at ICLR 2027

Enabling Token-wise Importance Sampling in Asynchronous RLVR with Proximal KL Regularization

Abstract

In asynchronous reinforcement learning with verifiable rewards (RLVR), rollout generation overlaps with policy optimization, so training responses may come from earlier policy versions. Policy-gradient methods commonly reuse these responses with token-wise importance weighting, avoiding products of likelihood ratios across long responses. Under appropriate conditions, we prove that full future KL credit preserves the exact proximal optimizer as the token-wise update's unique stationary policy. This holds despite gradient bias, even when responses come from a mixture of older policies. The update requires neither full-response importance weights nor a fitted critic. We further establish convergence of tabular continuous-time population updates under persistent coverage and time-varying behavior policies. These results motivate **Proximal KL**, a practical asynchronous RLVR recipe that separates the proximal reference from the sampling policy and uses bounded future KL credit in place of ordinary PPO clipping. We evaluate **Proximal KL** on Qwen3-1.7B and Qwen3-4B, finding improvements over our DAPO control: higher off-policy MATH500 and MinervaMath accuracy at both tested learning rates for 1.7B and higher AIME24 and AIME25 mean@16 and pass@16 at 4B. Ablations show better preservation of AIME mean@16 as rollout lag increases and reduced MATH500 sensitivity to baseline perturbations. The horizon sweep favors a finite future window over both tested extremes, while fixed-policy measurements show higher correction variance at longer tested horizons.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.