acceptodds
Under review as a conference paper at ICLR 2027

RhoNorm: Learning Channel-Energy Weights in Transformer Normalization

Abstract

In Transformers, shared normalization lets one channel's energy affect every channel's scale. RMSNorm fixes this aggregation rule to uniform weights. We learn the rule under ordinary predictive training and evaluate prediction and sparse-spike tolerance. Building on weighted divisive normalization, RhoNorm learns an exponentially parameterized positive diagonal metric, initialized to RMSNorm with one vector per site. Weighted energy shares link its task gradients to the exact shared-radius response under finite spikes. Across four decoder sizes, the learned rule improves clean prediction and reduces spike-induced extra loss in middle and late layers. In a 2.82B-parameter model, a spike in 1% of block-input coordinates at 75% depth adds 15.0 to the loss with RMSNorm and 6.3 with RhoNorm. Controls support relative energy weighting beyond output gains; the diagonal metric approaches dense-metric mean quality at much lower parameter cost. Quality gains extend to image classification and multimodal fine-tuning. On nanochat's time-to-GPT-2 target, a 1.38B-parameter model exceeds the GPT-2 reference in mean full CORE after 88 minutes of training on eight H100 GPUs, compared with 99 minutes for the fastest published Autoresearch entry.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.