Beyond Overconfidence: Logprob-Calibrated Advantage for LLM Reasoning
Abstract
The scaling of Large Language Model (LLM) reasoning via Reinforcement Learning is bottlenecked by prohibitive computational cost and reliance on external Process Reward Models (PRMs). While RM-free self-alignment via Group Relative Policy Optimization (GRPO) offers an efficient alternative, mining valid supervision from sparse outcomes faces three challenges: (1) Confidence-Blind Penalization, where advantage functions ignore internal generative certainty, treating overconfident hallucinations and uncertain guesses with identical severity; (2) Signal Dilution in Long Trajectories, where sparse global rewards fail to resolve credit assignment, masking localized logic fractures within reasoning chains; and (3) Calibration Divergence, where iterative policy updates sharpen output distributions without grounding them in factual truth, destroying the model's intrinsic self-awareness. To address these, we introduce LC-GRPO (Logprob-Calibrated Group Relative Policy Optimization), framing RM-free self-alignment as endogenous confidence calibration. We derive a Step-Level Generative Confidence metric directly from the policy's internal log-probabilities as a zero-cost pseudo-PRM, and introduce a Confidence-Decayed Advantage mechanism to apply non-linear penalties targeting early-onset, high-confidence reasoning errors. Experiments show LC-GRPO delivers a robust, well-calibrated self-evolution path, mitigating compounding errors and achieving superior reasoning efficiency while preserving strict alignment between model confidence and logical accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.