acceptodds
Under review as a conference paper at ICLR 2027

The Heavy-Tail Behavior of Adam: Sharp Rates and a Phase Transition at

Abstract

Gradient noise in deep learning produces more extremes than models with light tails predict; finite -moment models permit infinite variance (), beyond classical theory. Despite Adam's observed advantage over SGD, whose degradation is partly attributed to heavy tails, analyses rely on Clip-SGD analogies or modified methods, leaving unmodified coordinatewise Adam's convergence a recurring open problem. The difficulty is that conditional second moment proxies control dependence from sample reuse at , but may fail below . We make progress toward resolving this problem under a classical finite horizon parameter setting. Specifically, we derive an explicit upper bound for every constant base stepsize using chord cancellation and an exact identity for an auxiliary iterate. These charge normalization bias and drift in adaptive stepsizes to realized second moment increments, avoiding noise variances and direct momentum recursion. For -smooth objectives bounded below, with fixed and adapted noise having finite conditional -moments () bounded by , with fixed, we select a stepsize yielding the expected average squared gradient norm bound . Lower bounds for over all constant stepsizes establish stepsize minimax sharpness up to logarithms, with a gap of one logarithm for . Thus uniform polynomial stationarity decay is achievable iff , recovering the polynomial exponent at .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.