acceptodds
Under review as a conference paper at ICLR 2027

Learning from Correctness Trends: Sample-Wise Historical Advantage Shaping for Mathematical Reasoning

Abstract

Reinforcement learning is becoming an important approach for improving the reasoning ability of large language models. Group Relative Policy Optimization (GRPO) avoids the training cost of an additional value network through group-wise relative rewards and has been widely used for reasoning model training. However,the standard GRPO constructs its optimization signal primarily from the current candidate group and does not retain how the same training sample has evolved across previous visits. Consequently, candidate groups with similar current relative rewards receive similar update strengths even when their sample-wise learning trajectories differ. In this paper, we propose TrendGRPO, a sample-wise historical advantage shaping method that enhances within-group comparison with cross-iteration information. TrendGRPO maintains a correctness history for each sample across repeated visits and uses the deviation between the current group accuracy and historical accuracy to asymmetrically reshape positive and negative normalized advantages. We evaluate TrendGRPO on multiple base models and mathematical reasoning benchmarks. Ablations and training diagnostics support historical sample-wise conditioning in advantage shaping.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.