PC-GRPO: Pairwise Confidence Group Relative Policy Optimization for Post-Training Reasoning Language Models
Abstract
Reinforcement learning with verifiable rewards (RLVR) is an effective paradigm for post-training reasoning large language models (LLMs), and Group Relative Policy Optimization (GRPO) is a widely used critic-free optimizer in this setting. However, standard GRPO relies on group-wise reward normalization for advantage estimation, which provides coarse credit assignment and is particularly limited under sparse binary rewards. As a result, trajectories with different degrees of partial correctness often receive indistinguishable supervision. We propose Pairwise Confidence Group Relative Policy Optimization (PC-GRPO), which improves GRPO with two components: a pairwise confidence-based advantage estimator and a fine-grained verifiable reward. The first replaces coarse group-normalized advantages with pairwise comparisons that capture relative reward differences and policy confidence. The second provides denser supervision for incorrect trajectories using step-level overlap with reference reasoning traces, without introducing learned reward models. PC-GRPO remains critic-free and preserves the verifiable nature of RLVR. Across mathematical, scientific, knowledge-intensive, and coding benchmarks, PC-GRPO consistently outperforms standard GRPO and other strong baselines across multiple model scales.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.