Rethinking Group-Relative Advantage Estimation in Policy Optimization for Text-to-Text Regression
Abstract
Fine-grained scalar rewards provide rich supervision for reinforcement learning with large language models and naturally support text-to-text regression, where numerical prediction error can directly define the reward. However, as model predictions improve, rewards within rollout groups become increasingly concentrated, a phenomenon we refer to as reward collapse. Existing group-relative advantage estimators either allow the learning signal to shrink with these reward differences or normalize away the extent of reward concentration. We introduce Min-Max GRPO (**MM-GRPO**), which normalizes centered rewards by their within-group range, maintaining a well-scaled learning signal while allowing relative reward diversity to determine each group's contribution to optimization. We additionally introduce Rank-based GRPO (**R-GRPO**), which discards reward magnitudes and constructs advantages from within-group ranks to verify the importance of maintaining a well-spread learning signal throughout training. Across two real-world text-to-text regression benchmarks and two model sizes, MM-GRPO achieves the strongest overall regression performance, while R-GRPO remains competitive despite using only reward ordering. Additional analyses show that MM-GRPO and R-GRPO's improvements extend to long-tail target values and multilingual out-of-distribution inputs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.