Learning from Correctness Trends: Sample-Wise Historical Advantage Shaping for Mathematical Reasoning
Abstract
Reinforcement learning is becoming an important approach for improving the reasoning ability of large language models. Group Relative Policy Optimization (GRPO) avoids the training cost of an additional value network through group-wise relative rewards and has been widely used for reasoning model training. However,the standard GRPO constructs its optimization signal primarily from the current candidate group and does not retain how the same training sample has evolved across previous visits. Consequently, candidate groups with similar current relative rewards receive similar update strengths even when their sample-wise learning trajectories differ. In this paper, we propose TrendGRPO, a sample-wise historical advantage shaping method that enhances within-group comparison with cross-iteration information. TrendGRPO maintains a correctness history for each sample across repeated visits and uses the deviation between the current group accuracy and historical accuracy to asymmetrically reshape positive and negative normalized advantages. We evaluate TrendGRPO on multiple base models and mathematical reasoning benchmarks. Ablations and training diagnostics support historical sample-wise conditioning in advantage shaping.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.