Understanding Reward Normalization in GRPO
Abstract
Group Relative Policy Optimization (GRPO) has become a widely used and effective method for post-training large language models, yet the role of its within-group reward normalization remains underexplored. We identify reward discreteness as a key structure and develop a finite-sample theory for a family of normalized policy-gradient estimators. For binary rewards, we further show that the expected finite-group estimator is exactly collinear with its population counterpart. Our analysis reveals a group-size law that controls finite-group accuracy and motivates an adaptive sampler that assigns more rollouts to unresolved low-variance prompts. We interpret GRPO normalization as prompt reweighting: it strengthens low-variance prompt contributions while controlling aggregate gradient noise and exposes a tradeoff under more aggressive normalization. Finally, we analyze Balanced Aggregation, an existing token-aggregation rule, and prove that for binary rewards it preserves the population policy-gradient direction and obeys an analogous group-size law. Experiments on mathematical reasoning tasks compare normalization schemes and evaluate adaptive sampling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.