acceptodds
Under review as a conference paper at ICLR 2027

MeqMuon: Matrix-Equilibrating Muon for LLM Pretraining

Abstract

The success of large language models (LLMs) has been accompanied by continued growth in model size and pretraining costs. Muon offers high accuracy and training efficiency in LLM pretraining. Recent work introduces row-wise normalization into Muon to balance update magnitudes and improve pretraining performance. However, row-wise normalization alone cannot accommodate different imbalance patterns in update matrices. In this paper, we propose an improved Muon optimizer, called matrix-equilibrating Muon (), for LLM pretraining. balances both row and column magnitudes through normalization that can be automatically tailored to different imbalance patterns without manual intervention. Moreover, eliminates the need to store AdamW's second-moment estimates, reducing optimizer-state memory usage. Empirical results demonstrate that achieves better convergence performance than AdamW, Muon, and other baselines in LLM pretraining.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.