Structured Reward Reconstruction for Group-Normalized Policy Learning
Abstract
Group-normalized policy updates couple the rewards of all sampled completions. This coupling makes selective verification difficult: an unknown reward changes the update weights even for verified answers. For tasks with one canonical correct answer, we show that the complete update has a finite representation with one state per distinct answer class and one none-correct state. The representation makes classical prediction-correction estimators applicable to this nonlinear target and supports an optimal projection on its observation channel. Public coefficient vectors guide query selection, and detached reconstruction weights implement the estimate through an ordinary language-model loss. Online calibration adapts the reference distribution under a statewise covariance constraint. Experiments on GSM8K use three models, five seeds and a common training recipe. At 400 updates, generation-matched comparisons show fewer new checks than full feedback under shared stopping and answer-cache rules. Half-batch full feedback offers a competitive lower-generation alternative. Fixed-history analyses distinguish query-design effects from prediction correction and statewise covariance control; cost models quantify when fewer checks can offset additional computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.