SignMuon: Communication-Efficient Distributed Muon for LLM Pretraining
Abstract
Full-precision gradient exchange can limit the scaling of data-parallel training. We propose SignMuon, a distributed optimizer that combines Muon's matrix update directions with signSGD's majority-vote communication. Each worker approximates the polar factor of its local momentum matrix and communicates only its entrywise signs. Workers aggregate these signs using int8 AllReduce or packed 1-bit AllGather and apply the same update, keeping the polar computation local. We give conditional stationarity bounds for the actual momentum-polar vote and distinguish polar-solver error from the loss introduced by signing. On large language models (LLMs) up to 1.5B parameters trained on eight GPUs, SignMuon achieves lower perplexity at matched token budgets than the tested sign-based optimizers, while also showing favorable log-perplexity trajectories in wall-clock comparisons. We further evaluate optimizer memory usage and distributed communication efficiency in multi-GPU training. Together, these results show that SignMuon combines Muon's local matrix transformation with highly compressed global communication, delivering strong token efficiency and lower log perplexity for distributed LLM pretraining.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.