Rethinking KL Regularization in RLHF: From Wrong Value Estimation to Correct Gradient Optimization
Abstract
Reinforcement Learning from Human Feedback (RLHF) often uses a Kullback–Leibler (KL) penalty to stabilize training and control drift from a reference policy. However, recent implementations sometimes choose KL estimator formulas based on Monte Carlo value-estimation criteria such as bias and variance, and then reuse these formulas with gradient-carrying policy parameters as optimization losses. We show that this transfer from value estimation to loss design is wrong in general: good KL value estimation does not imply a correct KL gradient. We introduce a gradient-centric framework that compares KL implementations by the scalar coefficient they induce on the policy score function, covering both detached k_n in reward terms and directly differentiated k_n as loss terms. Through this framework, we first prove that k_1 as loss, although an unbiased KL value estimator, has zero expected on-policy gradient and provides no reference-directed regularization. We then prove that the PPO-style k_1 in reward implementation and the decoupled k_2 as loss implementation are gradient-equivalent and recover the Reverse KL (RKL) gradient under on-policy sampling. This equivalence, first proven here, identifies both as theoretically sound implementations of the RKL objective. In contrast, the GRPO-style k_3 as loss objective induces a biased first-order approximation to the principled RKL coefficient, with mismatched tail behavior. We further show that off-policy “as loss” KL terms require explicit importance-sampling corrections. Controlled experiments validate these findings, providing a comprehensive, gradient-based rationale for choosing and correctly implementing KL regularization in RLHF.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.