Analysis of the Stability of Preconditioned Stochastic Gradient Descent
Abstract
We study the algorithmic stability of preconditioned gradient descent (PGD) for empirical risk minimization. Although preconditioning affects finite-time stability, existing analyses fail to explain PGD's long-term behavior. We propose a refined stability analysis that separates the effect of preconditioning on transient optimization dynamics from PGD's intrinsic stability, recovering the long-term behavior obscured by standard contraction-based analyses. For deterministic updates, our decomposition establishes non-asymptotic bounds, while showing that long-term stability is independent of the preconditioner. For stochastic updates, we establish an analogous decomposition with an additional term that retains a dependence on the preconditioner. In both cases, each contribution is of order , up to variance terms and transient terms that decay with the number of iterations. Our conclusions hold in two regimes. First, the regularized setting, where individual losses are Lipschitz and the empirical objective is regularized by a sample-independent strongly convex function. Second, projected PGD for constrained problems, showing that projection introduces only additional residual terms. Finally, we use these stability estimates to derive expected generalization bounds and excess-risk guarantees for regularized and constrained empirical risk minimization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.