acceptodds
Under review as a conference paper at ICLR 2027

AdaSC-Muon: Adaptive Subspace Compression for Communication-Efficient Muon Training

Abstract

Muon has emerged as a promising optimizer for large-scale neural network training, but communication and orthogonalization costs can limit its efficiency in distributed settings. EF21-Muon by Gruntkowska et al. reduces communication by compressing momentum differences. Its best-performing compressor still requires workers to transmit local singular bases at every iteration, so the worker-to-server communication cost scales with the original matrix dimensions. We introduce adaptive subspace-compressed Muon (AdaSC-Muon), which uses shared subspaces for both momentum compression and Muon computation. The key insight is that aggregating projected momentum before orthogonalization enables a global update to be computed directly in a reduced matrix space. Workers reuse the shared bases and quantize projected momentum differences with error feedback. Periodic full-momentum corrections update the subspaces as momentum changes during training. We establish convergence guarantees and communication complexity bounds for AdaSC-Muon. Language model pretraining experiments with NanoGPT on FineWeb show that, compared with EF21-Muon, AdaSC-Muon and its diagonal variant reduce worker-to-server communication cost by up to 47.8% and reach the target validation loss in up to 53.8% less wall-clock time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.