Lower Bounds and Proximally Anchored SGD for Non-Convex Minimization Under Unbounded Variance
Abstract
Analysis of Stochastic Gradient Descent and its variants typically relies on the assumption of uniformly bounded variance, a condition that frequently fails in practical non-convex settings, such as neural network training, and in several elementary optimization settings. While several relaxations are explored in the literature, the classical Blum-Gladyshev (BG-0) condition, which permits the variance to grow quadratically with distance, has recently been shown to be the weakest condition. However, the study of the oracle complexity of stochastic first-order non-convex optimization under BG-0 has remained underexplored. In this paper, we address this gap and establish information-theoretic lower bounds, proving that finding an -stationary point requires stochastic BG-0 oracle queries for smooth functions and queries under mean-square smoothness. These limits demonstrate an unavoidable degradation from classical bounded-variance complexities, i.e., and for smooth and mean-square smooth cases, respectively. To match these lower bounds, we consider Proximally Anchored STochastic Approximation (PASTA), a unified algorithmic framework whose problem-dependent instantiations combine dynamic batching and gradient estimation with Halpern anchoring and Tikhonov regularization. We prove that suitable PASTA instantiations achieve minimax optimal complexities across numerous non-convex regimes, including standard smooth, mean-square smooth, weakly convex, star-convex, and Polyak-Lojasiewicz functions, entirely under an unbounded domain and unbounded stochastic gradients.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.