The Heavy-Tail Behavior of Adam: Sharp Rates and a Phase Transition at
Abstract
Gradient noise in deep learning produces more extremes than models with light tails predict; finite -moment models permit infinite variance (), beyond classical theory. Despite Adam's observed advantage over SGD, whose degradation is partly attributed to heavy tails, analyses rely on Clip-SGD analogies or modified methods, leaving unmodified coordinatewise Adam's convergence a recurring open problem. The difficulty is that conditional second moment proxies control dependence from sample reuse at , but may fail below . We make progress toward resolving this problem under a classical finite horizon parameter setting. Specifically, we derive an explicit upper bound for every constant base stepsize using chord cancellation and an exact identity for an auxiliary iterate. These charge normalization bias and drift in adaptive stepsizes to realized second moment increments, avoiding noise variances and direct momentum recursion. For -smooth objectives bounded below, with fixed and adapted noise having finite conditional -moments () bounded by , with fixed, we select a stepsize yielding the expected average squared gradient norm bound . Lower bounds for over all constant stepsizes establish stepsize minimax sharpness up to logarithms, with a gap of one logarithm for . Thus uniform polynomial stationarity decay is achievable iff , recovering the polynomial exponent at .
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.