FMuon: Coordinated Momentum Orthogonalization for Federated Training of Large Models
Abstract
Existing Federated Learning (FL) methods still primarily use element-wise local optimizers (e.g. AdamW, SGD) in each client, neglecting the geometric structure of the weight matrices. This often leads to the amplification of pathological directions in the weights during local updates, leading to deterioration in the condition number and slow convergence. Therefore, we introduce the Muon optimizer in each client (named Local Muon), which has matrix orthogonalization to optimize matrix-structured parameters. Experimental results show that, in the IID setting, Local Muon significantly accelerates the convergence of FL and reduces communication rounds compared to Local SGD and Local AdamW. However, in the non-IID setting, independent matrix orthogonalization based on the local distributions of each client induces strong client drift. Applying Muon in non-IID FL poses significant challenges: (1) Local–global lmo mismatch leading to client drift; (2) Momentum reinitialization. To address these challenges, we propose a novel Federated Muon optimizer (FMuon), which incorporates two key techniques: (1) momentum aggregation, where clients use the aggregated momentum for local initialization; (2) local-global alignment, where the local gradients are aligned with the global update direction to significantly reduce client drift. Theoretically, we establish a convergence guarantee for FMuon. Empirically, we validate the effectiveness of FMuon on language and vision models. Compared to several baselines including Local Muon, FMuon significantly reduces communication rounds and improves test accuracy. The code is available in https://anonymous.4open.science/r/FedMuon-935D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.