acceptodds
Under review as a conference paper at ICLR 2027

Where Does RLVR Act? A Conserved Token Ledger for What Post-Training Adds to a Base Model

Abstract

Most token-level accounts of what reinforcement learning with verifiable rewards (RLVR) changes in a base model are statistics of the two policies' next-token distributions, which have no total in accuracy and cannot say what a token is worth. Prefix transplantation with one permanent handoff (the trained policy writes tokens, the base the rest) is a conserved ledger: by the performance-difference lemma its increments are the trained tokens' expected advantages under the base's own value function and sum to the measured gain; a dense check at the first four positions finds no over-count on any cell and a small pooled under-count (+0.014), as action-set truncation predicts. Across 13 cells (0.5B–8B), where front-loading is largest (0.73 of the gain in four tokens), the policy's most frequent opening, grafted into every problem, reproduces the whole early gain, and a better template on the base supplies about half (46–66%) of the early lift of every Qwen-Math RLVR descendant tested: the early credit is mostly format, relative to the base's template. GRPO on a random reward delivers 0.80 of its gain in four tokens, all of it format. Each advantage is a room, set by the token's surprisal under the base (its rate, in nats), times a worth, its value over the base's alternatives. Divergence statistics track the room; none predicts the worth's size beyond weakly (rank correlation at most 0.26 in magnitude). Code, raw records and scripts regenerating every table are provided as supplementary material.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.