ADAPTRLHF: ADAPTIVE KL-CONSTRAINED REIN FORCEMENT LEARNING FROM HUMAN FEEDBACK WITH GROUP-NORMALIZED AD VANTAGE ESTIMATION
Abstract
Reinforcement Learning from Human Feedback (RLHF)has emerged as the dom inant paradigm for aligning large language models (LLMs) with human prefer- ences. However, existing RLHF methods suffer from two critical failure modes: (i) static KL penalty coefficients that cannot adapt to non-stationary training dynamics, causing either reward hacking or excessive conservatism; and (ii) high-variance advantage estimates due to unbounded cross-group reward hetero geneity. We propose AdaptRLHF, a novel RLHF framework that addresses both limitations through: (1) an adaptive KL controller that continuously ad justs the divergence penalty βt based on a proportional-integral (PI) feedback law over the observed KL gap; and (2) Group-Normalized Advantage Estima tion (GNAE) that standardizes rewards within semantically coherent generation groups, substantially reducing gradient variance. We further derive a unified objective that subsumes PPO-style clipping and GRPO-style group rewards un der a single theoretical framework. Extensive experiments on five reasoning benchmarks—GSM8K, ARC-Challenge, HumanEval, MMLU, and BIG-Bench Hard—demonstrate that AdaptRLHF achieves state-of-the-art results, outper forming PPO-RLHF and GRPO by +4 . 5% and +2 . 6% on average accuracy, with a measured 30% reduction in policy-gradient variance (Sec. 5.4) and 1.8 × faster convergence. Code and models are anonymously released at https: //anonymous.4open.science/r/AdaptRLHF.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.