acceptodds
Under review as a conference paper at ICLR 2027

Provably Finding Low Neural Loss via Spectral-GD

Abstract

Optimizers of the Muon family are rapidly becoming a competitive alternative to the standard adaptive methods for training LLMs. Such optimizers update the training parameters exploiting the intrinsically rich geometry of the space of the architectural weight matrices. In this work we do a first-of-its-kind analysis of Spectral-GD - the full-batch, momentum-free core of Muon - to show that it can provably train a shallow ReLU net in the large width regime. In terms of any set accuracy target and failure probability over random initialization, we give explicit estimates for the required width, step-length and number of iterations to reach the accuracy target. Our work rigorously establishes the surprise that gradient direction information, in the matrix sense, can be enough to train wide nets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.