acceptodds
Under review as a conference paper at ICLR 2027

Advantage Harness: Regulating GRPO's Advantage Signal by Shaping What It Sees

Abstract

Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each roll- out a magnitude from within-group reward statistics. This magnitude does not tell a reasoned correct answer from a guessed one, and it is largest when few rollouts in the group are correct. We identify this component as the spurious advantage. It arises on bounded-answer tasks and on the bounded sub-cases that cover more than half of MATH-7.5K. On multiple-choice prompts, about half of the correct rollouts in single-correct groups lack a supporting derivation, and training reduces this share but does not remove it. A zero-mean reweighting that gives one value to correct and one to wrong rollouts only rescales the group and cannot act on groups whose rollouts are all correct or all wrong. The lever that remains is the prompt distribution, which GRPO keeps fixed while the pass rate of the policy, and with it the set of groups that carry a gradient, moves during training. To this end, we propose Advantage Harness, a prompt selection layer that over-draws prompts at each update and keeps those least likely to yield a zero-gradient group. The esti- mator and the training rollout budget remain those of GRPO. On Qwen2.5-0.5B, the harness halves the zero-gradient rate and improves the eight-benchmark av- erage over GRPO, a shuffled control and a difficulty-label filter on every seed. On Qwen2.5-3B, a prior rebuilt from a checkpoint after 40 updates restores the zero-gradient reduction. Code will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.