acceptodds
Under review as a conference paper at ICLR 2027

LEARNING WHERE TO CONTINUE: BUDGETED CORRECTION OF CENSORED POLICY GRADIENTS

Abstract

In language-model training, reinforcement learning with verifiable rewards improves reasoning through outcome-based feedback. Generation limits leave some trajectories unfinished, while completing every rollout is costly. Imputation provides low-cost updates but leaves open where further generation most reduces gradient estimation error. We propose a budgeted continuation framework that selects rollout groups using joint prefix–tail residual risk and corrects their imputed updates using selected completions. Inverse-probability weighting over all arriving groups accounts for selective feedback when learning this risk. We establish unbiasedness for the unclipped, finite-horizon gradient of a frozen policy, derive the optimal allocation under an expected budget constraint, and bound the excess risk of its learned counterpart. An exact autoregressive construction shows that prefix–tail coupling can make allocation based on segment-wise second moments strictly suboptimal, even with a perfectly calibrated imputation head. Across six benchmarks in mathematics, coding, knowledge and long-context reasoning, joint-residual allocation increases mean accuracy by 6.27 and 6.37 percentage points on Qwen3.6-27B and Qwen3.6-35B-A3B, respectively, and reduces gradient correction error by 24.67% and 28.02% relative to fixed continuation at the same training budget. In ablations, accounting for prefix–tail coupling and correcting selective feedback each yields lower correction error and higher mean accuracy on both dense and mixture-of-experts models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.