RADeMuon: Provable Resource-Aware Quantized Decentralized Muon via Finite-Step Newton–Schulz Orthogonalization
Abstract
Distributed Muon optimizer enables matrix-aware optimization across multiple devices, but its full-precision local training and model exchange impose substantial resource costs, particularly when resource availability changes throughout decentralized training. We propose RADeMuon, a resource-aware quantized framework for decentralized Muon that supports precision configurations varying with available resources. RADeMuon performs local training with quantized parameters, gradients, and momentum states, applies finite-step Newton–Schulz (NS) orthogonalization to the quantized momentum, and directly communicates the resulting quantized parameter for decentralized aggregation. Theoretically, we trace quantization errors from their accumulation in the momentum state to their transformation by finite-step NS and subsequent propagation through decentralized aggregation. This analysis establishes a convergence guarantee whose dominant stationarity term is under vanishing-error conditions. Experiments on image classification and language-model pretraining show that RADeMuon reduces local computation and communication costs by up to 58.34% and 58.67%, respectively, relative to full-precision De-EF21-Muon while maintaining comparable model quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.