TraceBand: Prefix Reweighting under Finite-Sample Uncertainty for RLVR
Abstract
In reinforcement learning with verifiable rewards (RLVR), sampled continuations can be used to estimate the values of intermediate reasoning prefixes, but finite sampling leaves these estimates uncertain. This uncertainty makes it unclear whether observed differences reflect true value differences and therefore justify reweighting the sampled prefixes. We show that directly reweighting sampled prefixes using these finite-sample estimates can decrease the sampled-prefix objective. We introduce TraceBand, which constructs simultaneous confidence bands for prefix values and reweights sampled prefixes to maximize worst-case objective improvement. The resulting maximin problem has a closed-form solution that leaves the sampled-prefix distribution unchanged whenever the intervals have a nonempty intersection, that is, whenever all prefixes may have the same value, in which case reweighting could only incur a cost without any gain. Whenever the intervals contain the true prefix values, this reweighting has nonnegative gain in the sampled-prefix objective relative to the unchanged prefix weights. Experiments across mathematical reasoning and code generation show that TraceBand improves over finite-sample plug-in reweighting and achieves strong performance relative to representative RLVR baselines across multiple models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.