Can Normalization Match Muon at Scale?
Abstract
Muon is increasingly replacing AdamW to train large-scale neural networks, but it requires expensive matrix orthogonalization. Recent studies report gains over Muon with simple row or column normalization alone. However, we find that normalization does not match Muon at scale. Our analysis shows that normalization leaves large correlations across rows, while Muon removes them. We also study how normalization should be combined with Muon. Applying normalization after Muon is more effective than applying it before, while using EMA statistics for normalization provides little benefit. The improvement over vanilla Muon decreases at larger scales. We establish these findings in the overtrained regime of frontier production language models. Our models are trained at 37–110 the Chinchilla optimum, whereas NanoGPT-style benchmarks commonly used in recent work are close to 1 Chinchilla. Experiments across vision tasks (e.g., supervised learning, self-supervised learning, diffusion models) lead to similar conclusions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.