On the Optimality of Adam under Heavy-Tailed Noise
Abstract
Adam is the canonical optimizer in deep learning, and its decoupled weight decay variant AdamW is widely used for training large language models (LLMs). Meanwhile, Muon has recently achieved strong empirical gains over AdamW. One proposed explanation is robustness to heavy-tailed gradient noise: Muon admits sharp convergence rates, whereas Adam could behave poorly in such settings. This contrast motivates a basic question of whether Adam can converge optimally under heavy-tailed noise. We answer affirmatively through a stronger result for Adam-mini, a blockwise framework that encompasses Adam and Adam-Norm. Under heavy-tailed th noise moment, we prove that Adam-mini attains the sharp rate . Our theory also supports bias correction and weight decay, thereby covering AdamW as a special case. We further establish a matching lower bound that proves minimax optimality and, as a consequence, yields the first general stochastic stationarity lower bound for smooth non-convex optimization with an unavoidable dependence, where is the problem dimension. Together, these results sharply characterize heavy-tailed convergence for Adam and its variants. Comparing the Adam directly with Muon under the same problem class, we show that, unlike prior justifications, heavy-tailed noise alone cannot explain Muon's empirical advantage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.