SONU: Communication-Efficient Orthonormalized Updates via Sketching
Abstract
As Large Language Model (LLM) training scales across distributed systems, gradient synchronization can incur substantial communication overhead, particularly in bandwidth-constrained settings. Gradient compression reduces this cost by communicating compact approximations of the gradient. Existing approaches, including sparsification and quantization, are typically designed to preserve the gradient under entrywise or Euclidean reconstruction criteria. These criteria, despite effective in coordinate-wise preconditioning optimizers, need not preserve the information used by matrix-orthonormalized optimizers such as Muon, whose update depends on the gradient’s spectral geometry rather than its individual entries. Such discrepancy makes preserving the geometry an ideal target of the gradient compressor. We propose Sketched Orthonormalized Updates (SONU), a communication-efficient optimizer built on two-sided gradient sketching that facilitates synchronizing gradients in a low-dimensional space. SONU forms a generalized Nyström surrogate of the gradient from the sketches, which coincides with the dense gradient up to the sketched subspaces, and extracts the paired singular vectors without dense orthonormalization. To prevent a fixed complement from being persistently omitted, SONU maintains an evolving sketch, in which the sketching subspace is updated with the orthonormalized directions and are gradually transported across iterations. Empirically, on chain-of-thought, agentic function-calling, and instruction-following tasks, SONU attains lower perplexity per communicated byte than the other communication-efficient orthonormalized update algorithms. Furthermore, SONU matches the perplexity trajectory of Dion, an efficient Muon variant which synchronizes the full gradients before applying orthonormalized updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.