RAC-GRPO: RISK-CONSTRAINED ADVANTAGE CALIBRATION FOR FINITE ROLLOUT GROUPS
Abstract
Group Relative Policy Optimization (GRPO) constructs advantages from reward statistics estimated within finite rollout groups. As a result, its training weights are sensitive to finite-group sampling. With binary rewards, this issue is most extreme when all sampled responses are correct or all are wrong, in which case GRPO assigns zero advantage to the entire group. Mixed groups are also affected because their reward means and standard deviations are estimated from only a few samples. We introduce Risk-Constrained Advantage Calibration for GRPO (RAC-GRPO), a stateless calibration algorithm based on the observed success count and group size. It adjusts the reward statistics by adding symmetric virtual success and failure counts. A reference posterior describes uncertainty in the underlying success probability. A mean squared error constraint limits the allowable correction strength. Within this constraint and a fixed cap, RAC-GRPO selects the largest admissible correction. The resulting advantages decompose into an attenuated within-group contrast and a shared offset. This recovers small updates for all-correct and all-wrong groups while also calibrating mixed groups. Under independent on-policy sampling with equally weighted responses, we characterize the expected update. We show that endpoint recovery is aligned with the success gradient. The recovered component’s mean and directional standard deviation both scale linearly with recovery magnitude. Training requires only a precomputed lookup table, with no additional responses, reward history, or learned critic. Across seven mathematical reasoning benchmarks, RAC-GRPO improves over GRPO by 2.53, 2.08, and 2.82 percentage points on Qwen3-1.7B, 4B, and 8B, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.