NORMALISATION CHOICES IN DECODER-ONLY TRANSFORMERS ARE REGIME-DEPENDENT
Abstract
Three claims about normalisation in Transformers circulate as folk theorems: Pre-LN is stable and Post-LN is not; RMSNorm is a free drop-in for LayerNorm; and QK-Norm stabilises attention. They are rarely tested against each other in one codebase, on one corpus, with shared seeds. We do so with 175 runs of a GPT-2-style decoder (21.3M non-embedding parameters at depth 12) on FineWeb-Edu, comparing six variants arranged as two single-axis chains rooted at Pre-LN and Post-LN, five of them across four depths, a learning-rate sweep, reruns at the selected rates, a gradient-clip stress test and per-layer gradient diagnostics. The apparent answer to “which variant is better?” changes materially with the optimisation regime. At a shared learning rate the Post-LN penalty is +0.0495 nats; at each variant’s sweep-selected learning rate it is +0.1442 nats, a 2.9-fold change from a single hyperparameter. A single-architecture chain attributes the Post-LN rescue cleanly: swapping LayerNorm for RMSNorm moves the loss by +0.0001 nats, while adding QK-Norm moves it by −0.0504 nats and lifts the Post-LN learning rate ceiling threefold. That headroom is depth-limited: at depth 48 five variants converge to within 0.021 nats, but at 10−3 both Post-LN variants—QK-Norm included—collapse whether gradient clipping is relaxed 100× or left at its default, so of the two variables crossed the learning rate determines the outcome and the clip does not. The failures are not blow-ups: collapsed runs settle +0.020 nats from the corpus unigram entropy, which we measure directly. We release all 175 per-run artefacts with a pipeline that regenerates every number here.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.