acceptodds
Under review as a conference paper at ICLR 2027

Incorrect Answers Can Be Reinforced More: Relative-Reinforcement Policy Optimization for Learning from Verifiable Rewards

Abstract

Group relative policy optimization (GRPO) trains reasoning models from verifiable rewards by favoring correct responses within each sampled group. However, we observe that this local preference does not always carry over to what training accumulates: measured against the initial policy, the incorrect responses to a question can be reinforced more than its correct ones. We call this phenomenon reinforcement inversion and attribute it to the local nature of GRPO's signal, which ranks responses within the current group but does not constrain how accumulated change is divided between correct and incorrect responses. To address this, we propose Relative-Reinforcement Policy Optimization (\rpo), which adds a group-level loss on the reinforcement gap between the two classes, weighted by class balance and shaped by a logistic function. \rpo reuses GRPO's rollouts, verifier labels and reference log-probabilities, requiring no additional sampling. Experiments on mathematical and multimodal reasoning tasks across two backbones show that \rpo consistently improves over GRPO and outperforms alternative policy optimization methods across models and datasets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.