acceptodds
Under review as a conference paper at ICLR 2027

STABLE ALIGNMENT: ADAPTIVE KL CONTROL WITH GROUP-NORMALIZED ADVANTAGE ESTIMATION FOR RLHF

Abstract

Reinforcement Learning from Human Feedback (RLHF) has become a standard approach for aligning large language models, yet policy-gradient training remains notoriously unstable: reward scale changes across batches, group-level statis tics are estimated from only a few samples, and KL divergence from the ref erence policy drifts unless the penalty coefficient β is carefully tuned. We in troduce Stable Alignment (SA), a lightweight modification of RLVR/PPO-style training that couples two mechanisms. First, G-NAE (Group-Normalized Ad vantage Estimation) centers advantages within each prompt group and rescales them using a pooled within-group variance estimate. Pooling removes between group difficulty shifts while reducing variance of the scale estimator by O(1/N) in batch size, avoiding both noisy per-group scaling and the between-group at tenuation of global normalization. Second, AKC (Adaptive KL Control) uses an incremental proportional–integral controller on log β so that empirical KL tracks a user-specified budget. Because G-NAE makes the policy update approx imately invariant to reward scale, the KL response to β becomes more station ary, which is precisely the regime in which fixed-gain feedback is reliable. On a controlled verifiable-reward benchmark, SA matches the final accuracy of GRPO and global-normalization baselines while providing penalty-coefficient robustness across three orders of magnitude of βinit and closed-loop KL-budget tracking. Empirically, the pooled scale estimator has about 3 × lower cross-batch coeffi cient of variation than the per-group estimator, supporting the theoretical variance reduction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.