acceptodds
Under review as a conference paper at ICLR 2027

Within-Class Score Variation in Standardized Group-Relative Policy Gradients

Abstract

Success rates determine how often a binary-reward group contains a learning signal, but they do not determine the variability of its policy-gradient update. For the on-policy, unclipped standardized estimator used in Group Relative Policy Optimization (GRPO), we derive an exact finite-group covariance decomposition into success-count variation and score variation within each reward class. A four-outcome policy shows that the latter can change while success probability and the expected update remain fixed. We also derive the estimator's influence function: estimating the reward standard deviation contributes to the leading covariance, even though that estimate is consistent. In an audit of 3,072 responses from two Qwen checkpoints, within-class variation accounts for median covariance fractions of 97.8% and 82.4% at group size eight, measured in fixed last-layer LoRA coordinates and conditional on mixed empirical response pools. Validation with 4,096 independent responses separates this algebraic identity from moment-estimation accuracy: 32-response calibration pools underpredict aggregate covariance by about 16% for the smaller model. A separate nine-run sampling-allocation study gives mean held-out accuracy of 57.67% for covariance-informed allocation and 57.11% for uniform allocation; a descriptive paired interval for the difference includes zero. Success rates therefore omit a substantial source of update variation in the measured score space, but using this additional information for sampling has yet to yield a reliable accuracy gain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.