acceptodds
Under review as a conference paper at ICLR 2027

From Correct Answers to Correct Supervision in Self-Rewarded Test-Time Learning

Abstract

Self-rewarded learning reinforces the answers that win a model's votes. A correct response can therefore be available without becoming useful supervision. We study this selection step through recorded reward groups, controlled oracle-reward mixtures, and diagnostics of adaptation outcomes. Across 74 self-rewarded Llama-3-8B runs on screened GSM8K, a correct answer appears in 91.5% of training groups but is selected in 63.7%. We derive the group-dependent oracle weight at which an available correct response acquires a positive centered reward; its median is 1/3 among wrong-vote groups containing that response. Across 288 reward-intervention runs, oracle anchoring raises mean held-out gain and reduces its variability. Selection correctness also provides information before the trajectory is complete: using the first 19 of 38 recorded groups, it reduces leave-one-seed-out gain-prediction error by 19.2% beyond response accuracy, with a 16.1% reduction in the reward-mixture cohort after controlling reward weight. These results distinguish producing a correct answer from selecting it for reinforcement, and establish reward selection as a measurable link between available responses and adaptation outcomes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.