acceptodds
Under review as a conference paper at ICLR 2027

How Much of the Natural Gradient's Value Do Cheap Approximations Retain?

Abstract

Cheap approximations of natural gradient descent (NGD), such as K-FAC, are justified by the argument that a direction close to the natural gradient inherits its benefits. We test this argument on small dense and convolutional networks whose exact Fisher is computable and strongly ill-conditioned. Every update is normalised to unit Fisher norm, damping and horizon are shared, and each method tunes its own step size. We score each method by its retention: the fraction of exact NGD's improvement over gradient descent that it achieves. Under this protocol, on one-hidden-layer networks with at most parameters, K-FAC, EK-FAC and the exact diagonal retain – of the training-loss improvement, and Kronecker approximations beat diagonal ones in every configuration tested. The angle to exact NGD ranks the approximations, but a GD-to-NGD interpolation at the same angle predicts only a third to a tenth of their retention. At a Euclidean angle, error in the lowest-curvature decile keeps all of NGD's benefit, while error in the highest-curvature decile makes the update far worse than gradient descent; at a matched damped-Fisher angle this contrast largely disappears, so it mainly reflects the Fisher-norm step budget. On dense networks, most retention survives at exact NGD's step size. These results point to relative curvature fidelity as a promising design target. In this regime, the exact layer-wise block retains , so much higher fidelity is achievable, at higher cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.