Adam under Bounded Variance: Polylogarithmic Confidence without Clipping
Abstract
High-probability guarantees quantify the reliability of Adam and other adaptive optimizers. Under bounded noise variance, jin2026sharper show that Adam's second moment normalization already yields polylogarithmic confidence for the preconditioned gradient energy, while their ordinary squared gradient guarantee retains polynomial dependence on . This motivates asking whether Adam can preserve that confidence benefit in ordinary stationarity. We show that adaptive normalization coupled with second moment memory of order achieves this without clipping or stronger noise assumptions. For smooth objectives bounded below and conditionally unbiased noise with bounded variance, we combine with a suitably chosen constant learning rate, any fixed first moment coefficient, and positive second moment initialization. For any fixed positive offset, , and , the average squared gradient norm is with probability at least . When the target confidence is specified before the run, adjusting the memory and learning rate sharpens the guarantee to under the same assumptions. Descent localized by objective height and a potential coupling accumulation with displacement control the cumulative gradient cost of large denominators and give both guarantees through a common analysis.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.