acceptodds
Under review as a conference paper at ICLR 2027

Prompt-Level Advantage Shaping and -GRPO

Abstract

Reinforcement learning methods such as PPO and its variants are widely used to improve pretrained LLMs. GRPO simplifies PPO by replacing the learned critic with group-based estimates, but can yield weak gradient signals for both simple and hard prompts. Practical variants therefore use heuristic techniques such as prompt-level advantage shaping to modify the contribution of individual prompts. While shaping can strengthen the learning signal, it also biases the gradient estimator, which may no longer converge to stationary points of the mean-reward objective. This paper contributes to the theoretical understanding of GRPO. Based on recent Bernstein polynomial computations for the optimization bias, we analyze the expected gradient change caused by advantage shaping. Explicit formulas allow to understand how different prompts contribute to the expected gradient. We introduce -GRPO, a family of GRPO variants with a possibly time-varying degree of prompt-level advantage shaping. By annealing this weighting over training, -GRPO can retain the learning signal provided by prompt-level shaping early in training while recovering convergence to stationary points of the mean-reward objective. The mathematical feasibility of -GRPO is shown using biased SGD arguments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.