acceptodds
Under review as a conference paper at ICLR 2027

Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities

Abstract

Reinforcement learning with verifiable rewards (RLVR) can improve reasoning accuracy while producing increasingly concentrated outputs. Under binary rewards, Group Relative Policy Optimization (GRPO) assigns identical positive advantages to correct responses within mixed-reward groups, without accounting for their generation confidence. We propose Probabilistic GRPO (ProGRPO), which introduces a confidence-based Advantage Re-weighting Mechanism (ARM). ARM uses the difference between prompt and response confidence, computed as geometric means of token probabilities over low-probability positions under the old policy. Within each reward class, lower-confidence responses receive larger adjusted advantage coefficients. The correction is disabled for uniform-reward groups and preserves advantage signs under a sufficient bound on its magnitude. Experiments on mathematical reasoning and code generation using Qwen2.5 and DeepSeek-R1-Distill-Qwen models show improvements in accuracy and multi-sample solution coverage. On Qwen2.5-7B, adding ARM to a matched GRPO with Clip-Higher baseline improves average Pass@1 and Pass@32 by 3.3 and 8.3 percentage points, respectively. The complete ProGRPO method improves these metrics over vanilla GRPO by 5.7 and 13.8 percentage points. Training-entropy and output-similarity measurements provide complementary evidence of reduced output concentration in the evaluated settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.