acceptodds
Under review as a conference paper at ICLR 2027

Provable Benefits of Block Preconditioning over Gradient Descent under Block-Heterogeneous Smoothness

Abstract

Motivated by the empirical advantage of adaptive optimizers on Transformers, we study the comparison between block-preconditioned gradient descent (Block-PGD) and constant-step gradient descent (GD). We introduce block-heterogeneous smoothness, in which parameter blocks have distinct smoothness constants Lₖ, and measure stationarity by G² = Σₖ ‖∇ₖf‖²/Lₖ, the ordinary gradient norm in curvature-normalized coordinates. Block-PGD attains min(t < T) G² ≤ 2Δ/T, whereas GD attains 2RₕΔ/T, where Rₕ is the ratio of the largest to the smallest block smoothness constant. A quadratic instance proves that linear dependence on Rₕ is necessary for GD, including under loss and ordinary-gradient criteria, although it does not match the general nonconvex dependence on accuracy. We extend the upper bounds to stochastic gradients, establish a local Zipfian softmax majorizer with ratio V raised to the power α, and give a conditional RMSProp-style result with √Rₕ dependence. Direct Hessian–vector-product diagnostics on a small GPT-2 model trained on WikiText-2 find persistent architecture-block curvature heterogeneity and strong second-moment/curvature alignment for output-class blocks. These results explain the benefit of block preconditioning; they do not constitute a convergence theory for full Adam with momentum.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.