Normalization removal at scale: RMSNorm-to-affine transition in GPT-OSS-style MoE transformers
Abstract
Although normalization is nearly universal in modern pre-norm Transformers, it is still unclear whether input-dependent normalization statistics are required by the learned computation or if they primarily serve to stabilize optimization. To answer this question, we examined three GPT-OSS-style mixture-of-experts (MoE) transformer variants with 1 billion, 3 billion, and 6 billion total parameters each. These models were trained on 8 billion, 16 billion, and 32 billion processed tokens, respectively, and evaluated on 8 common NLP benchmarks. Our main intervention is a homotopy, or convex combination, from RMSNorm to a scalar affine map at each layer. This combination is parameterized by a coefficient that decreases from 1 (pure RMSNorm) to 0 (pure affine function) during training. The optimal slope and intercept of the affine function were chosen to minimize the mean squared error between RMSNorm and its best linear approximation. Across scales, the normalization-free endpoints retain most of the zero-shot downstream performance of their normalized baselines. Nevertheless, the performance gap becomes more pronounced at larger scales. We also trained learnable element-wise alternatives from step zero. Among the trained variants, the variant initialized as tanh(20x) achieved the best performance, though it remained behind the RMSNorm baseline. The results demonstrate that a simple affine endpoint and a carefully scheduled gate can eliminate the need for normalization statistics. However, the results also demonstrate that the schedule's effectiveness diminishes as the scale increases. Altogether, these results suggest treating normalization as primarily an optimization-stability mechanism and searching for richer, self-normalizing non-linearities that can replace its stabilizing role without the need for explicit, sample-dependent normalization statistics. The code and checkpoints to reproduce this study are available at: URL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.