From Correct Answers to Correct Supervision in Self-Rewarded Test-Time Learning
Abstract
Self-rewarded learning reinforces the answers that win a model's votes. A correct response can therefore be available without becoming useful supervision. We study this selection step through recorded reward groups, controlled oracle-reward mixtures, and diagnostics of adaptation outcomes. Across 74 self-rewarded Llama-3-8B runs on screened GSM8K, a correct answer appears in 91.5% of training groups but is selected in 63.7%. We derive the group-dependent oracle weight at which an available correct response acquires a positive centered reward; its median is 1/3 among wrong-vote groups containing that response. Across 288 reward-intervention runs, oracle anchoring raises mean held-out gain and reduces its variability. Selection correctness also provides information before the trajectory is complete: using the first 19 of 38 recorded groups, it reduces leave-one-seed-out gain-prediction error by 19.2% beyond response accuracy, with a 16.1% reduction in the reward-mixture cohort after controlling reward weight. These results distinguish producing a correct answer from selecting it for reinforcement, and establish reward selection as a measurable link between available responses and adaptation outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.