Proper Scoring under Group Reward Normalization: Target Distortion and Exact Finite-Group Updates
Abstract
GRPO trains language models from several answers to one prompt and standardizes their rewards to stabilize learning. For probability answers scored by one shared random outcome, prior work found overconfidence with two proper scores, which favor honest reports, in a large-group approximation, but left open whether this failure is general or persists in finite groups. We identify the key mechanism: the shared outcome determines both the rewards and their normalization scale, which can turn a small win and a large loss into equally strong updates. We precisely characterize the favored probability with group statistics fixed and prove that every such binary reward fails in some setting; even the best available answer can reverse. We derive the exact update when each sampled answer also changes its own denominator. Centering without division preserves the mean gradient direction at fixed group size; an independent scale shared across outcomes preserves each prompt's target. Qwen2.5-3B experiments connect exact updates to neural gradients and show higher scoring loss with standardization than with centering alone under plain SGD and complementary budget controls.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.