acceptodds
Under review as a conference paper at ICLR 2027

TraceBand: Prefix Reweighting under Finite-Sample Uncertainty for RLVR

Abstract

In reinforcement learning with verifiable rewards (RLVR), sampled continuations can be used to estimate the values of intermediate reasoning prefixes, but finite sampling leaves these estimates uncertain. This uncertainty makes it unclear whether observed differences reflect true value differences and therefore justify reweighting the sampled prefixes. We show that directly reweighting sampled prefixes using these finite-sample estimates can decrease the sampled-prefix objective. We introduce TraceBand, which constructs simultaneous confidence bands for prefix values and reweights sampled prefixes to maximize worst-case objective improvement. The resulting maximin problem has a closed-form solution that leaves the sampled-prefix distribution unchanged whenever the intervals have a nonempty intersection, that is, whenever all prefixes may have the same value, in which case reweighting could only incur a cost without any gain. Whenever the intervals contain the true prefix values, this reweighting has nonnegative gain in the sampled-prefix objective relative to the unchanged prefix weights. Experiments across mathematical reasoning and code generation show that TraceBand improves over finite-sample plug-in reweighting and achieves strong performance relative to representative RLVR baselines across multiple models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.