AdaSC-Muon: Adaptive Subspace Compression for Communication-Efficient Muon Training
Abstract
Muon has emerged as a promising optimizer for large-scale neural network training, but communication and orthogonalization costs can limit its efficiency in distributed settings. EF21-Muon by Gruntkowska et al. reduces communication by compressing momentum differences. Its best-performing compressor still requires workers to transmit local singular bases at every iteration, so the worker-to-server communication cost scales with the original matrix dimensions. We introduce adaptive subspace-compressed Muon (AdaSC-Muon), which uses shared subspaces for both momentum compression and Muon computation. The key insight is that aggregating projected momentum before orthogonalization enables a global update to be computed directly in a reduced matrix space. Workers reuse the shared bases and quantize projected momentum differences with error feedback. Periodic full-momentum corrections update the subspaces as momentum changes during training. We establish convergence guarantees and communication complexity bounds for AdaSC-Muon. Language model pretraining experiments with NanoGPT on FineWeb show that, compared with EF21-Muon, AdaSC-Muon and its diagonal variant reduce worker-to-server communication cost by up to 47.8% and reach the target validation loss in up to 53.8% less wall-clock time.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.