What Do Rollout Tokens Buy? Hidden Costs and Budgeted Accuracy in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) pays for every token it generates, yet training choices are rarely judged by what those tokens buy. We study this question on competition mathematics with two training choices and one evaluation choice. First, we replace one third of a SmolLM3-3B training set with harder problems and keep the other 2,048 problems at the same positions. Responses to these unchanged problems become twice as long (+101%) at a similar training reward, and late in training most of their tokens are in responses that reach the length limit. Second, we halve the training response limit of Qwen3.5-9B from 4,096 to 2,048 tokens. Training compute falls by 28%, and majority-vote accuracy under a 16,384-token budget rises from 36.1% to 43.5% as far more evaluation answers finish. The gain remains when truncated answers are continued to 8,192 tokens. Third, under the same budget, counting problems with any correct answer and counting the answer a model returns can rank a base model and its RL-trained version in opposite orders. The SmolLM3-3B base finds a correct answer on more of 216 problems (57 against 39) but returns a correct majority answer on fewer (5 against 18), and majority vote favors the trained model in all three model pairs. In all three cases, more tokens did not mean more correct answers. We therefore suggest judging an RLVR training choice by the tokens it adds on problems that stay the same and by the answers the trained model returns under a fixed token budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.