Exact Tail Laws for Nonlinear SGD with Infinite-Variance Gradients
Abstract
With infinite-variance gradient noise, SGD can have infinite expected excess objective even when most iterates are close to the optimum. We characterize its accuracy through exact first-order tail probabilities for the last iterate of nonlinear, constant-stepsize SGD. For regularly varying innovations of index , the leading measure transports one large gradient shock through the local Hessian flow. This is the same measure that governs the central stable Ornstein–Uhlenbeck limit; we prove its relative accuracy at shrinking tolerances above the central scale, with state-dependent noise and a growing observation horizon. The admissible window is nonempty for every available gradient moment . The main technical step is a moment bound at the central fluctuation scale and a lower-moment control of the nonlinear error. The result yields explicit projection and excess-objective tails, extreme quantiles, and a factor for fixed batch size . For affine population gradients, a sharper comparison gives the same law at fixed tolerances, including random sample Hessians. Coupled simulations in dimensions one to one hundred examine relative calibration, objective quantiles, batch effects, and the error due to random curvature.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.