acceptodds
Under review as a conference paper at ICLR 2027

Pretraining Data Scale Reverses the Muon-AdamW Fine-Tuning Gap

Abstract

Muon has recently emerged as a promising optimizer for pretraining large language models, offering a speedup over AdamW in minimizing pretraining loss. However, it remains unclear whether Muon’s advantage persists after fine-tuning. In this paper, we study the downstream behavior of Muon and AdamW and uncover a surprising trend: beyond a sufficiently large pretraining data budget, AdamW outperforms Muon after fine-tuning. Moreover, as model size increases, AdamW overtakes Muon at progressively smaller pretraining token–parameter ratios. We observe the same trend under Gaussian perturbations of the pretrained weights, where Muon-pretrained models degrade increasingly more than AdamW-pretrained models as the data budget grows. We complement our empirical findings with a controlled study of Muon and AdamW in a linear model. In this setting, we show theoretically and empirically that the margin-to-noise ratio, a tractable proxy for robustness motivated by a second-order approximation of the loss, shifts from favoring Muon to favoring gradient descent as the data budget grows, mirroring the fine-tuning trend.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.