Normalize your normalized layers
Abstract
Normalization, both in the forward pass and in the optimizer, is a central component of modern deep learning. But where and how should we normalize? We study this question from a purely optimization-based perspective and find that the answer prescribes a placement of normalization layers consistent with state-of-the-art architectures. Exploiting the scale invariance induced by normalization layers, we develop NorSODA, a weight-normalized method that achieves noise-robust accelerated rates by carefully incorporating weight decay. Our key technical result is an online-to-batch conversion that wraps any no-regret algorithm while only making assumptions on the sphere. Interestingly, the result extends to non-Euclidean geometries, yielding a close cousin of MuonH that removes the need to normalize the update direction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.