STABLE ALIGNMENT: ADAPTIVE KL CONTROL WITH GROUP-NORMALIZED ADVANTAGE ESTIMATION FOR RLHF
Abstract
Reinforcement Learning from Human Feedback (RLHF) has become a standard approach for aligning large language models, yet policy-gradient training remains notoriously unstable: reward scale changes across batches, group-level statis tics are estimated from only a few samples, and KL divergence from the ref erence policy drifts unless the penalty coefficient β is carefully tuned. We in troduce Stable Alignment (SA), a lightweight modification of RLVR/PPO-style training that couples two mechanisms. First, G-NAE (Group-Normalized Ad vantage Estimation) centers advantages within each prompt group and rescales them using a pooled within-group variance estimate. Pooling removes between group difficulty shifts while reducing variance of the scale estimator by O(1/N) in batch size, avoiding both noisy per-group scaling and the between-group at tenuation of global normalization. Second, AKC (Adaptive KL Control) uses an incremental proportional–integral controller on log β so that empirical KL tracks a user-specified budget. Because G-NAE makes the policy update approx imately invariant to reward scale, the KL response to β becomes more station ary, which is precisely the regime in which fixed-gain feedback is reliable. On a controlled verifiable-reward benchmark, SA matches the final accuracy of GRPO and global-normalization baselines while providing penalty-coefficient robustness across three orders of magnitude of βinit and closed-loop KL-budget tracking. Empirically, the pooled scale estimator has about 3 × lower cross-batch coeffi cient of variation than the per-group estimator, supporting the theoretical variance reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.