acceptodds
Under review as a conference paper at ICLR 2027

SR-GRPO: Reinforcement Learning with Self-Refinement from Verifiable Rewards

Abstract

Reinforcement learning from verifiable rewards (RLVR) with a rule-based binary outcome verifier produces strong chain-of-thought reasoners, yet it rewards only the final answer. The policy never learns to inspect or correct its own intermediate steps, so any error in the reasoning chain propagates to the final answer. Addressing this self-refinement gap faces a dilemma: inference-time prompting methods are deployable but, without ground-truth feedback, often degrade accuracy; training-time approaches do learn to self-refine, but mostly by using a supervision signal stronger than RLVR's binary verifier—teacher correction traces, step-level annotations, or separately trained process reward models. We introduce SR-GRPO (Group Relative Policy Optimization with Self-Refinement), a two-stage RLVR framework in which a single shared policy acts as both solver (generating an initial solution) and refiner (correcting its own mistakes), with both stages trained jointly by GRPO. SR-GRPO uses only the binary verifier RLVR already requires—no external teacher, no separate critic, and no ground-truth answer at refinement. SR-GRPO has three mechanisms: wrong-first sampling keeps self-refinement focused on genuine solver errors; a ground-truth-free refinement prompt makes the refinement stage deployable at inference, since the model never sees the answer when it refines; and self-distilling feedback recycles every verified refinement into the solve stage as additional supervision, so the same policy serves as solver, critic, and teacher within a single training loop. Under a single fixed recipe across three backbone families and two model scales (1.5B and 4B), SR-GRPO matches or surpasses the strongest RL baselines on several math reasoning benchmarks: at 1.5B it matches QwQ-32B with far fewer parameters; at 4B, where vanilla GRPO saturates, SR-GRPO continues to improve, surpassing several open-source 20B+ reasoners. Beyond the math benchmarks, SR-GRPO also improves on out-of-distribution general-reasoning benchmarks (GPQA Diamond, MMLU-Pro) across all three backbone families—evidence that the gains are not math-specific reward hacking. Double-blind LLM judges prefer SR-GRPO's reasoning over strong RL baselines, and qualitative analysis confirms that SR-GRPO genuinely locates and revises erroneous steps during self-refinement. Overall, SR-GRPO turns self-refinement into a training-time self-distillation procedure that adds no inference-time overhead and generalizes across backbone families, model scales, and reasoning domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.