acceptodds
Under review as a conference paper at ICLR 2027

AdaRLHF :Stabilizing RLHF via PID KL Control and LOO Group Advantages

Abstract

Reinforcement learning from human feedback (RLHF)must balance reward improvement against divergence from the reference policy, yet a static KL penalty βeither invites reward hacking or over-constrains exploration,and group normalized advantages of the kind popularized by GRPO remain imprecisely characterized. We present AdaRLHF,a framework that addresses both issues with an emphasis on verifiable analysis.First,we recast KL coefficient tuning as a discrete-time tracking problem and derive an explicit proportional –integral – derivative (PI+D)law with anti-windup and an annealed divergence budget,for which we prove a non-stationary tracking-error bound under mild Schur condi tions (Prop.1).Second,we replace GRPO-style advantages with a self-free,leave one-out group-normalized estimator (LOO-GNAE)and give an exact variance characterization under exchangeable,intra-group correlated rewards (Thm.1), showing that (i)the full-group baseline introduces a standardization bias propor tional to the reward skewness,(ii)the leave-one-out baseline is immune,and (iii) the variance premium of self-free baselines is exactly (G/(G _1))2 ,independent of the correlation . The bias–variance trade-off admits a closed-form interpola tion λ★ (Cor.1). Our analysis also corrects a variance formula reported in prior adaptive-RLHF work.Experiments on six reasoning benchmarks with five seeds show that AdaRLHF improves average accuracy over GRPO by 6.8 points and PPO-RLHF by 9.1 points,reduces advantage variance by 42%,and eliminates the late-training reward-hacking collapse;component ablations attribute the gains and validate each theoretical prediction empirically.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.