MuonMix: A Mixed-LMO Block-Periodic Algorithm for Efficient Distributed Muon Training
Abstract
The Muon optimizer accelerates language model pretraining by applying Newton-Schulz orthogonalization to hidden weight matrices, but scaling it to distributed training requires gathering gradient shards across devices at every step. MuonBP reduces this cost through block-periodic orthogonalization, but applying the spectral-norm linear minimization oracle (LMO) to local shards can introduce a partition-dependent approximation error. Since max-norm LMO updates do not suffer from this error, we propose MuonMix, a mixed-LMO block-periodic algorithm for distributed Muon training that assigns each hidden weight matrix to either the spectral-norm LMO or the max-norm LMO under a shared synchronization period , based on a per-layer selection heuristic motivated by a layerwise convergence upper bound. Layers assigned to the max-norm LMO require no gather or scatter operations and introduce no block approximation error. We evaluate the proposed algorithm on language model pretraining at 160M, 960M, and 1.2B parameters on FineWeb, and show that it improves throughput by 3.0%, 4.8%, and 6.5% over MuonBP and achieves faster early convergence across all model sizes, and reaches target validation losses up to 7.1% faster in wall-clock time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.