acceptodds
Under review as a conference paper at ICLR 2027

Parallel Layer Normalization for Universal Approximation

Abstract

Layer Normalization (LN) is not only an optimization component: between affine layers it also supplies nonlinearity, and grouping changes how this nonlinearity scales with width. We study this mechanism through universal approximation. In Parallel Layer Normalization networks (PLN-Nets), fixed-dimensional LN units are applied independently to groups between two affine layers. We prove that a shallow LN-Net with one LN unit is non-universal regardless of width, whereas PLN-Nets become universal as the number of parallel groups grows. The mechanism extends to deep networks, Root Mean Square Normalization (RMSNorm) variants, position-wise feed-forward networks, and convolutional networks under suitable grouping structures. Controlled MLP and physics-informed neural network (PINN) experiments support PLN's practical viability, while CNN and ViT experiments show that group count shapes the finite-sample fitting envelope but does not act alone: the grouped axes, replacement site, and label structure lead to distinct fitting and generalization behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.