acceptodds
Under review as a conference paper at ICLR 2027

SignMuon: Communication-Efficient Distributed Muon for LLM Pretraining

Abstract

Full-precision gradient exchange can limit the scaling of data-parallel training. We propose SignMuon, a distributed optimizer that combines Muon's matrix update directions with signSGD's majority-vote communication. Each worker approximates the polar factor of its local momentum matrix and communicates only its entrywise signs. Workers aggregate these signs using int8 AllReduce or packed 1-bit AllGather and apply the same update, keeping the polar computation local. We give conditional stationarity bounds for the actual momentum-polar vote and distinguish polar-solver error from the loss introduced by signing. On large language models (LLMs) up to 1.5B parameters trained on eight GPUs, SignMuon achieves lower perplexity at matched token budgets than the tested sign-based optimizers, while also showing favorable log-perplexity trajectories in wall-clock comparisons. We further evaluate optimizer memory usage and distributed communication efficiency in multi-GPU training. Together, these results show that SignMuon combines Muon's local matrix transformation with highly compressed global communication, delivering strong token efficiency and lower log perplexity for distributed LLM pretraining.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.