Activation-Adjacent Layer Normalization: Effects on Approximation Power
Abstract
Layer normalization (LN) and root mean square normalization (RMSNorm) remain input-dependent and nonlinear at inference, so inserting either next to an activation need not preserve represented functions. We study when activation-normalized (AN) and preactivation-normalized (PN) shallow networks can approximate the function represented by a prescribed width- one-hidden-layer MLP. Viewing RMSNorm as input-dependent scaling, we use width augmentation to locally linearize normalization on compact feature sets. For AN networks, an unbounded activation range allows width , while every continuous nonconstant activation admits width with error . For PN networks, affine parameters or, without them, degree-one positive homogeneity recover width ; general continuous activations admit width with error tending to zero. The corresponding LN constructions use centered coordinate pairs, yielding widths and , respectively. We prove that for RMSNorm and for LN are also the worst-case bound for approximations, and show that applicable universal-approximation theorems for ordinary MLPs transfer to these normalized architectures. Numerical experiments recover the predicted approximation rates for both normalizations across all five settings and illustrate the shifted-ReLU finite-width obstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.