Correct Reward, Costly Rule: Gradient Coherence Decides What Group-Relative RL Locks In
Abstract
When a verifiable reward is exact but the policy's observation is insufficient, the reward-maximising policy must follow a feature that is right only on average, and its cost falls entirely on the samples that feature misleads. We study how group-relative RL reaches this outcome, using next-POI ranking as a controlled testbed. First, GRPO converges to the constrained-optimal rule: it locks onto the most readable feature, lifts aggregate accuracy, pushes the misled subgroup below chance, and keeps applying the rule where the feature is uninformative. Second, the cost is asymmetric because aligned gradients cohere while misaligned ones cancel, not because either side is larger. Third, an exact advantage–position identity shows that this coherence is created by centring the advantage, not inherited from the model's representations. Fourth, the resulting criterion predicts which feature is locked in before training, even against the base model's preference. Finally, re-anchoring the advantage moves the cost exactly as the diagnosis implies, and bounds what any reweighting of sampled rollouts can recover.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.