acceptodds
Under review as a conference paper at ICLR 2027

StructMuon: Compact Core Communication by Reusing Muon History Subspaces

Abstract

Muon improves large model training through momentum orthogonalization, but scaling requires efficient communication across devices. Low-rank compression can reduce communication costs, but existing methods compute their projection bases separately from the optimizer and must transmit these bases or recompute them with SVDs. Can the optimizer itself provide these bases? We find that Muon's momentum accumulates gradient history, while its orthogonalized update retains paired left and right singular directions. We extract history subspaces from these shared optimizer states, and each worker can reproduce them locally. We propose StructMuon, which uses these subspaces as communication bases and transmits a compact core. A residual channel exploits concentrated residual energy to send information outside these subspaces, while local error feedback retains unsent information. With contractive residual compression, we prove communication error contraction without SVD safeguards and derive a nonconvex stationarity bound that tightens with better subspace capture. Experiments on GPT-2 Large and NanoGPT-124M show 92.0-97.5% less communication than Full Muon synchronized once per update at matched validation perplexity. These results suggest that communication can reuse the optimizer's own structure.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.