Pretraining Data Scale Reverses the Muon-AdamW Fine-Tuning Gap
Abstract
Muon has recently emerged as a promising optimizer for pretraining large language models, offering a speedup over AdamW in minimizing pretraining loss. However, it remains unclear whether Muon’s advantage persists after fine-tuning. In this paper, we study the downstream behavior of Muon and AdamW and uncover a surprising trend: beyond a sufficiently large pretraining data budget, AdamW outperforms Muon after fine-tuning. Moreover, as model size increases, AdamW overtakes Muon at progressively smaller pretraining token–parameter ratios. We observe the same trend under Gaussian perturbations of the pretrained weights, where Muon-pretrained models degrade increasingly more than AdamW-pretrained models as the data budget grows. We complement our empirical findings with a controlled study of Muon and AdamW in a linear model. In this setting, we show theoretically and empirically that the margin-to-noise ratio, a tractable proxy for robustness motivated by a second-order approximation of the loss, shifts from favoring Muon to favoring gradient descent as the data budget grows, mirroring the fine-tuning trend.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.