Future-KL Regularized GRPO: Process-Level Credit Assignment from -Divergence Regularization
Abstract
GRPO is widely used for critic-free LLM post-training, but its KL regularization is usually implemented as a local loss-side token penalty. We show that this misses the policy-gradient signal induced by autoregressive KL regularization. Unlike standard KL-regularized RL objectives, GRPO's group normalization induces a non-linear prompt-level utility; for binary verifier rewards, this utility is . As a result, reward and KL cannot be fused before normalization without changing the implicit objective. We derive the on-policy gradient of GRPO-style objectives with token-wise -divergence regularization. The reward term recovers the standardized GRPO advantage, while the regularizer contains a causal future cost-to-go omitted by local KL losses. For reverse KL, this yields Future-KL Regularized Policy Optimization (FRPO), a critic-free correction implemented by a reverse cumulative sum of sampled token log-ratios after advantage construction. We further study a windowed variant that interpolates between local and full-horizon regularization credit. On mathematical-reasoning tasks, full-horizon FRPO improves pass@16 in the experiments, maintaining a better exploration ability compared to other regularization methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.