acceptodds
Under review as a conference paper at ICLR 2027

Denoising-GRPO: Critic-Free Confidence-Gated Consensus Reward Correction for Noisy Reasoning

Abstract

Group Relative Policy Optimization (GRPO) is a practical critic-free method for reasoning reinforcement learning (RL), yet it struggles with noisy supervision, such as inconsistencies between prompts and expected labels. Such noise compromises group-relative advantage estimation, slowing convergence and, in the worst case, steering the model toward degenerate solutions. To address this bottleneck, we propose Denoising-GRPO, a highly practical and scalable drop-in reward-side extension of GRPO. Denoising-GRPO converts within-group answer agreement into a structured correction target. To prevent the policy from being overridden by unstable or incorrect majorities during early training, it uses the model's internal generation confidence to apply this correction only when the group is deemed reliable. Because it operates entirely on the reward construction side, our method avoids heavier auxiliary estimators or external reward models and leaves the efficient critic-free GRPO outer loop unchanged. To thoroughly benchmark this capability, we construct noisy training variants of textual MATH and multimodal Geo3K, simulating realistic corruption scenarios during supervision. All reported test and benchmark results are evaluated on clean held-out data with correct ground-truth targets. We evaluate Denoising-GRPO after both low-noise and high-noise training, and test transfer on GSM8K, held-out same-corpus MATH500, Minerva-Math, MathVista, MathVerse, and MMK12. On clean in-domain tests after high-noise training, Denoising-GRPO improves over GRPO by up to +10.5 on MATH and +3.0 on Geo3K, and also improves clean benchmark performance (e.g., +10.1 on GSM8K). Mechanism diagnostics further show that the confidence gate reduces harmful consensus overrides while recovering most of the gap to clean-supervision GRPO. These results indicate that reasoning models can self-denoise corrupted training supervision and still generalize to clean evaluation data whenever within-group agreement is sufficiently reliable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.