acceptodds
Under review as a conference paper at ICLR 2027

Policy Lag in GRPO: Group-Estimator Bias, Optimizer Memory, and Delayed Feedback

Abstract

Policy lag makes a learner optimize older-policy responses while rollout reward evaluates that older policy. We diagnose this mismatch in group relative policy optimization (GRPO) through its group-gradient estimator, optimizer memory, and feedback channel. In a ten-arm study with 20 paired seed blocks and a 16-update refresh interval, token-level truncated importance sampling (TIS) reduces confirmed terminal rollout failures from 12/20 to 2/20 and raises mean final-learner accuracy from 22.7% to 57.9%. Selective resets identify the first moment as the strongest reset target, and local norm-matched controls favor its update direction. We derive the cross-response bias that survives response-level importance correction and use exact enumeration to separate directional alignment, estimator bias, and finite-batch error. In a disjoint monitoring cohort with 30 development and 50 held-out runs, current-learner evaluation gives 28/35 timely alerts versus 18/35 from rollout reward, with one false alert among 13 no-episode runs for each. These findings distinguish correcting stale-data updates, managing optimizer history, and observing current-policy quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.