acceptodds
Under review as a conference paper at ICLR 2027

From Curvature Fidelity to Step Reliability: What GGN Proxies Capture and Miss

Abstract

Curvature-aware optimizers train neural networks by using a local loss model to guide the direction and size of parameter updates. They often replace the Hessian with the generalized Gauss–Newton (GGN) matrix, a positive-semidefinite surrogate that supports practical second-order computation. But when does a more faithful curvature model lead to a more reliable update, especially when the update is evaluated on a different batch? We propose a diagnostic framework linking curvature comparisons to predicted and realized loss changes. Within a common measured subspace, operator-pair geometry distinguishes differences in curvature directions from differences in normalized spectral shape. An exact error accounting then separates the effect of replacing the resolved GGN from the remaining Hessian contribution, finite-step nonlinearity, and evaluation-batch transfer. Automatic-differentiation products and dense reference checks support these measurements. Across independent ResNet and vision-transformer image-classification cohorts, transfer between batches dominates the measured absolute-error budget, largely through gradient differences. The GGN and resolved positive-Hessian models select identical updates from the original candidates; matching their predicted first-order decrease makes curvature differences affect selection. The primary loss comparison remains inconclusive, while secondary evaluation targets favor the GGN selector. These results explain why curvature fidelity, influence on step selection, and realized loss reduction must be assessed together, providing a basis for diagnosing which part of a local model matters for an update.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.