acceptodds
Under review as a conference paper at ICLR 2027

HFC-GRPO: Strict-Lag Selective Advantage Estimation for Zero-Variance Reward Groups

Abstract

Group Relative Policy Optimization (GRPO) loses its reward-derived advantage signal when all responses to a prompt receive the same binary reward. We introduce HFC-GRPO, a strict-lag selective advantage estimator that recovers signed advantages for these homogeneous groups while preserving GRPO advantages exactly on mixed groups. The estimator separates where to intervene from how much: the observed group type determines intervention and sign, while a historical forecast formed before generation determines the homogeneous-group correction magnitude. This design uses historical prediction to complement, rather than replace, within-group comparison. Our analysis identifies the distinction between all-rollout calibration and calibration on the outcome-selected homogeneous population, and bounds the resulting advantage and gradient-contribution errors under explicit assumptions. Across two Qwen3-8B reasoning protocols, paired comparisons favor HFC-GRPO, including a 0.65-percentage-point Acc@8 gain over RL-ZVP under a common training and evaluation pipeline. Transfer to multi-turn search agents yields a 2.12-percentage-point mean exact-match gain over paired outcome-only GRPO. Targeted replacements examine the roles of historical alignment, selective intervention, and adaptive prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.