Accelerating Stochastic Schatten Momentum Descent Against Unbounded Gradient Noise
Abstract
Matrix optimizers which exploit the spectral geometry are increasingly getting adopted for training large transformer based models. Inspired by these successes we seek to build foundations for these novel ML algorithms. Thus we are led to study a notion of stochastic Schatten momentum descent with Exponential Moving Average (EMA) and Momentum Variance Reduction (MVR). We introduce a generalized notion of smoothness and an affinely growing bound on the variance of stochastic gradient differences, both defined in the Schatten norm. Under these relaxed conditions of gradient smoothness and unbounded gradient noise, we prove (a) an optimal convergence rate for the variance reduced form of our algorithm by exploiting the aforementioned control on stochastic gradient differences and (b) a rate for the EMA variant. We further extend the EMA guarantee to Polyak's heavy-ball momentum, and thus getting new guarantees on the idealized stochastic Muon algorithm.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.