acceptodds
Under review as a conference paper at ICLR 2027

Can Normalization Match Muon at Scale?

Abstract

Muon is increasingly replacing AdamW to train large-scale neural networks, but it requires expensive matrix orthogonalization. Recent studies report gains over Muon with simple row or column normalization alone. However, we find that normalization does not match Muon at scale. Our analysis shows that normalization leaves large correlations across rows, while Muon removes them. We also study how normalization should be combined with Muon. Applying normalization after Muon is more effective than applying it before, while using EMA statistics for normalization provides little benefit. The improvement over vanilla Muon decreases at larger scales. We establish these findings in the overtrained regime of frontier production language models. Our models are trained at 37–110 the Chinchilla optimum, whereas NanoGPT-style benchmarks commonly used in recent work are close to 1 Chinchilla. Experiments across vision tasks (e.g., supervised learning, self-supervised learning, diffusion models) lead to similar conclusions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.