RhoNorm: Learning Channel-Energy Weights in Transformer Normalization
Abstract
In Transformers, shared normalization lets one channel's energy affect every channel's scale. RMSNorm fixes this aggregation rule to uniform weights. We learn the rule under ordinary predictive training and evaluate prediction and sparse-spike tolerance. Building on weighted divisive normalization, RhoNorm learns an exponentially parameterized positive diagonal metric, initialized to RMSNorm with one vector per site. Weighted energy shares link its task gradients to the exact shared-radius response under finite spikes. Across four decoder sizes, the learned rule improves clean prediction and reduces spike-induced extra loss in middle and late layers. In a 2.82B-parameter model, a spike in 1% of block-input coordinates at 75% depth adds 15.0 to the loss with RMSNorm and 6.3 with RhoNorm. Controls support relative energy weighting beyond output gains; the diagonal metric approaches dense-metric mean quality at much lower parameter cost. Quality gains extend to image classification and multimodal fine-tuning. On nanochat's time-to-GPT-2 target, a 1.38B-parameter model exceeds the GPT-2 reference in mean full CORE after 88 minutes of training on eight H100 GPUs, compared with 99 minutes for the fastest published Autoresearch entry.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.