The Variance Cost of Partial Reward Verification in Group-Based Policy Optimization
Abstract
Partial reward verification can preserve a complete-label finite-batch surrogate gradient in expectation while increasing its error at a fixed verification cost. We study when the information purchased by an audit compensates for its cost, taking the fully verified finite-batch surrogate gradient as the target. Using established unbiased estimators, we express this comparison as a threshold on residual-gradient reduction and evaluate it with direct measurements of all 7.37 million trainable LoRA coordinates. A fixed two-stage rule has 2.20–2.77 times the cost-matched whole-group MSE on four BIRD development batches: its first audit creates gradient error in groups whose complete reward vectors are constant. A separate 128-task test shows that cost accounting can reverse the ranking of a frozen whole-group allocation, from an MSE ratio of 0.720 at SQL-body CPU to 1.076 at complete verifier-child CPU. Finally, arithmetic-risk predictors improve development comparisons but do not confirm that gain on 64 tasks from eight additional databases. Together, these fixed-gradient experiments identify residual reduction, the execution cost boundary, and budget transfer as distinct requirements for economical verification.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.