acceptodds
Under review as a conference paper at ICLR 2027

V-Trace Policy Optimization: When Multi-Step Correction Helps, and Why

Abstract

Reinforcement learning for language models increasingly departs from the on-policy ideal: rollouts are reused across optimizer passes, actors lag learners, and inference and training engines disagree on token probabilities. PPO and GRPO correct only the current token, whereas exact trajectory importance sampling corrects the whole sequence at prohibitive variance. We show that prefix, current-action, and suffix ratios reweight the visited history, the sampled action, and its continuation, respectively. We introduce V-Trace Policy Optimization (V-TRACE PO), which attaches a coefficient-truncated suffix trace directly to the per-token score. One reverse scan computes the trace, and on complete terminal-reward responses it requires no learned critic; an exact identity says which continuation policy the truncated trace targets, and whatever the truncation, the correction adds no bias to the per-token score. We evaluate it in two families of individually preregistered experiments: one in which the same rollouts are reused across several optimizer passes, and one in which every sample is fresh but the behavior policy differs from the learner by construction—a lagging generator, a hotter sampler, or sequential minibatches— together with a sweep of the truncation cap, so the correction is also tested at full strength. The change-of-measure argument that motivates the estimator applies to the second family and not the first. The results come out the other way round: V-TRACE PO helps under reuse, and reliably in none of our fresh-sample controls, where widening the correction is if anything harmful. We analyze both halves. On fresh samples the correction does not pay: at every truncation we test, the variance it adds exceeds the bias it removes — even at the mildest mismatch we can construct, the numerical disagreement between the inference and training engines. Under reuse it helps for a reason the construction does not predict: the trace silently becomes an asymmetric learning rule, amplifying whatever an earlier pass over the same batch already did — promoting what succeeded and damping what failed. Under heavy sample reuse the improvement of V-TRACE PO over GRPO is substantial, and it comes with a wider usable trust region — hence less clipping — where that same widening degrades GRPO

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.