acceptodds
Under review as a conference paper at ICLR 2027

Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise

Abstract

Preconditioned methods are central to deep learning optimization. Predominant approaches include computationally light diagonal preconditioners such as Adam, which rely on gradient statistics, and second-order methods such as Gauss-Newton (GN), which capture richer curvature information. Seeking the best of both worlds, we disentangle the preconditioner design space into several factors, separating the choice of diagonal scaling (Adam-style versus GN-style) from 1) the _basis choices_ under which the diagonal scaling operates, and 2) the _gradient noises_ from mini-batching. Our theoretical results show that GN's optimality for linear regression no longer holds under a poorly chosen basis, high gradient noises, or a move from linear regression to a non-convex variant of logistic regression. In these settings, Adam-style methods offer genuine advantages over curvature-inspired preconditioning, rather than serving merely as a tractable proxy. Empirical results on synthetic problems and CIFAR-10 support these findings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.