acceptodds
Under review as a conference paper at ICLR 2027

Correct Reward, Costly Rule: Gradient Coherence Decides What Group-Relative RL Locks In

Abstract

When a verifiable reward is exact but the policy's observation is insufficient, the reward-maximising policy must follow a feature that is right only on average, and its cost falls entirely on the samples that feature misleads. We study how group-relative RL reaches this outcome, using next-POI ranking as a controlled testbed. First, GRPO converges to the constrained-optimal rule: it locks onto the most readable feature, lifts aggregate accuracy, pushes the misled subgroup below chance, and keeps applying the rule where the feature is uninformative. Second, the cost is asymmetric because aligned gradients cohere while misaligned ones cancel, not because either side is larger. Third, an exact advantage–position identity shows that this coherence is created by centring the advantage, not inherited from the model's representations. Fourth, the resulting criterion predicts which feature is locked in before training, even against the base model's preference. Finally, re-anchoring the advantage moves the cost exactly as the diagnosis implies, and bounds what any reweighting of sampled rollouts can recover.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.