Matched Compute Is Not Matched Signal: Where the Gradient Goes in Group-Relative Reinforcement Learning
Abstract
Reinforcement learning from verifiable rewards is usually preceded by a supervised warm-up, and runs are compared at matched compute. But group-relative methods such as GRPO learn only from prompts whose sampled answers disagree: a group of rollouts that all earn the same reward contributes exactly zero gradient, however much it cost to generate. Two runs with the same budget can therefore receive very different amounts of learning signal. Prior work uses this to filter prompts during training; we make it exact and extend it to policies and prompt sets. We derive the gradient signal a group carries under GRPO, Dr.GRPO and RLOO and average it over the prompts. The result is a signal efficiency , where is mean accuracy, is a prompt's success probability and depends only on the group size. A more accurate policy can therefore carry less signal. Across 32 runs given identical budgets, the generations that actually carried a gradient differed by a factor of . A warm-up trained on the problems its reinforcement-learning phase will use silences them: after epochs, solution distillation solves of them and falls from to . The same accounting explains why a base model barely improves: its gradient is real, but almost all of it comes from the format reward rather than from being right. Choosing prompts by , from a single sampling pass before training, beats a random set of the same size and budget on all runs across four settings, a search task and GSM8K at 3B to 14B, by to points of held-out accuracy. The lead survives at twice the training horizon, though its signal advantage fades as the chosen prompts are solved, and reverses on every seed in the one setting where the random control climbs far enough toward . Two plausible variants of the rule each fail in some setting, by overshooting an optimum the accounting locates. What does not do is rank warm-ups or predict how much a run will gain: it prices prompts, not checkpoints. Runs should therefore be compared by the generations that carry a gradient, not the generations spent. Code: https://anonymous.4open.science/r/signal-efficiency-rl-D5C1.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.