acceptodds
Under review as a conference paper at ICLR 2027

Progressive Reasoning Reinforcement Learning via Confidence-Aware Reward Reweighting

Abstract

In this paper, we propose Progressive Reasoning Reinforcement Learning (ProRL) to improve the reasoning capabilities of large language models. Existing RLVR methods, such as Group Relative Policy Optimization (GRPO), rely on binary outcome rewards that treat trajectories with the same outcome indiscriminately, overlooking their varying reliability and learning value. We argue that sampled trajectories exhibit different levels of mastery and should contribute unequally to policy optimization. Based on this insight, we instantiate ProRL with a likelihood-aware reward reweighting strategy that uses the likelihood under the historical policy, detached from gradient computation, as an estimate of trajectory confidence to dynamically adjust trajectory rewards. This enables progressive consolidation of reliable reasoning, refinement of emerging capabilities, and correction of dominant errors while preserving reasoning diversity. Experiments with Qwen2.5-1.5B on MATH-lighteval and DAPO-MATH-17K, together with scaling experiments using Qwen2.5-7B on MATH-lighteval, demonstrate that ProRL consistently outperforms strong RLVR baselines on both Pass@1 and Pass@16, improving reasoning accuracy while preserving diversity in solution exploration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.