Momentum Makes (Adaptive) Subgradient Methods Optimal for Stochastic Nonsmooth Nonconvex Optimization
Abstract
SGDM (stochastic gradient descent with momentum) consistently outperforms vanilla SGD across a wide spectrum of modern deep learning tasks, including the training of residual networks and Transformer-based large-scale models. Yet, theoretical understanding of this performance advantage remains incomplete, especially for the nonsmooth and nonconvex settings that dominate real-world deep learning applications. From an online learning perspective, a line of recent studies cutkosky2023optimal,ji2026derandomized,zhang2024random has established that carefully modified SGDM variants attain optimal convergence rates for locating Goldstein-style stationary points, thereby providing theoretical support for the empirical benefits of momentum acceleration. Despite these fruitful advances, existing studies heavily depend on additional algorithmic modifications such as model randomization or momentum clipping that deviate notably from practical SGDM implementations. In this paper, we pursue an alternative approach to directly analyze the convergence of standard SGDM over a broad family of weakly convex optimization problems. Equipped with squared-geometric-CDF randomized output, we prove that SGDM achieves optimal convergence rates under a mildly relaxed notion of Goldstein stationarity. Our results are valid for constant, decaying, and adaptive learning rate schedules alike. Furthermore, as a secondary contribution, our analysis extends to a slightly modified SGDM with one-time momentum reset, yielding novel theoretical insights into its recently observed empirical performance in training large language models and deep reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.