Incorrect Answers Can Be Reinforced More: Relative-Reinforcement Policy Optimization for Learning from Verifiable Rewards
Abstract
Group relative policy optimization (GRPO) trains reasoning models from verifiable rewards by favoring correct responses within each sampled group. However, we observe that this local preference does not always carry over to what training accumulates: measured against the initial policy, the incorrect responses to a question can be reinforced more than its correct ones. We call this phenomenon reinforcement inversion and attribute it to the local nature of GRPO's signal, which ranks responses within the current group but does not constrain how accumulated change is divided between correct and incorrect responses. To address this, we propose Relative-Reinforcement Policy Optimization (\rpo), which adds a group-level loss on the reinforcement gap between the two classes, weighted by class balance and shaped by a logistic function. \rpo reuses GRPO's rollouts, verifier labels and reference log-probabilities, requiring no additional sampling. Experiments on mathematical and multimodal reasoning tasks across two backbones show that \rpo consistently improves over GRPO and outperforms alternative policy optimization methods across models and datasets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.