Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon
Abstract
Muon accelerates language-model pretraining, but why its uniform spectral scaling works and where it falls short remain unclear. We probe singular directions of momentum buffers at training checkpoints and estimate their per-rank loss-optimal step sizes on held-out batches. Across models and training stages, we observe a stable spectral profile. The loss-optimal step size is relatively flat across the step-size-tolerant bulk but drops sharply at the leading direction, termed the step-size-sensitive head. This profile provides a spectral-allocation account of the performance ordering from SGD through Adam to Muon. It also suggests that Muon still underuses the tolerant bulk. We introduce Spectral-Aware Muon (SAMuon), which holds the head at the same scale as Muon while increasing the scale of the bulk using a static spectral prior. SAMuon follows the measured profile, whereas SAMuon-lite uses a simpler two-level approximation. Neither variant adds persistent optimiser state. Under idealised exact whitening, both variants retain the asymptotic convergence rates of Muon. Across evaluated language-model settings with 124M to 1B parameters, both variants achieve lower final validation loss than tuned AdamW and Muon (Scion implementation). SAMuon reaches the same validation loss as Muon using 13.3% to 24.0% fewer training tokens. SAMuon-lite retains most of this improvement and runs at nearly the same wall-clock speed as Muon in the measured 1B-model setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.