acceptodds
Under review as a conference paper at ICLR 2027

Grothendieck Meets Adam: One Probe, Three Budgets in Adaptive Optimization

Abstract

Sign descent and Adam work in the same geometry but pay for curvature differently: the standard smoothness that governs normalized descent and the adaptive smoothness that governs adaptive methods can differ by a factor of the dimension. We show that these two constants, and a pointwise constant between them, are one quantity: the largest expected curvature that a random probe can detect. They differ only in how the probe's budget is enforced: surely, on average at each point, or on average over the landscape. For convex objectives in diagonal geometry, adaptive smoothness is a squared -summing norm and the ideal preconditioner is its Pietsch measure. Each relaxation is governed by one invariant of the preconditioner algebra. Averaging at a point is a semidefinite relaxation whose gap lies between the largest matrix-block size and ; in diagonal geometry it is at most Grothendieck's constant , or for convex objectives, both asymptotically sharp. Averaging over the landscape costs a curvature rotation number, which is at most the rank of the algebra, the number of directions a preconditioner weighs independently, a bound that convex objectives approach; the rank also replaces the dimension in the known comparison of the two smoothness constants. Rotation is one for quadratics and binary logistic regression and grows linearly for log-sum-exp. In box-constrained convex optimization, rotation is the acceleration horizon: up to one logarithm, the minimax risk of deterministic first-order methods over functions of rotation at most is for queries and for , and one adaptive method, analysed without choosing a basis, attains both rates in every well-structured geometry. Gradient noise obeys the same three levels.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.