Spectral Momentum Steering Mitigates Optimizer Mismatch of Muon Finetuning
Abstract
Muon has emerged as a compelling alternative to Adam for training matrix-valued parameters and has been adopted in the pretraining of frontier language models. Its use in post-training, however, is hindered by optimizer mismatch: fine-tuning an Adam-pretrained checkpoint with Muon can yield worse performance than fine-tuning with Adam itself. Because most publicly available checkpoints are pretrained with Adam, resolving this mismatch would enable broader application of Muon to post-training. Analyzing how pretrained singular directions align with layer representations, we find that pretrained weights read from input directions associated with spectral outliers beyond both edges of the Marchenko–Pastur (MP) distribution, yet write predominantly into the subspace spanned by the principal left singular vectors. Building on this, we propose Spectral Momentum Steering (SMS), which performs steepest descent under a plasticity-weighted anisotropic norm constraint that limits movement along the principal output directions encoding foundational representation structure. Across Adam-pretrained model families, SMS consistently improves pretrained capability preservation while matching or exceeding the downstream adaptation performance of matched Adam or standard Muon fine-tuning, substantially alleviating optimizer mismatch.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.