acceptodds
Under review as a conference paper at ICLR 2027

A mechanistic theory of feature learning

Abstract

We present a theory of feature learning in overparameterized neural networks built on the dynamics of the “active parameter subspace”: the low-dimensional subspace of parameter space spanned by the network Jacobian. We show that this subspace is determined by the instantaneous gradient of the network with respect to its parameters, that it evolves throughout training, and that the curvature of the loss acts as the infinitesimal generator of its evolution. Thus, feature learning becomes a trajectory on the Grassmannian of parameter subspaces, driven by the Hessian of the loss, which in turn depends on weight decay. This perspective contrasts with the neural tangent kernel (NTK) approximation, in which learning reduces to coefficient updates within a fixed basis, and with phenomenological kinematics accounts that describe feature evolution without an underlying generator. Numerical experiments on fully connected networks trained with full-batch gradient descent support the picture: weight decay sustains the subspace trajectory after training loss has plateaued, enabling the network to discover a representation that also generalizes, while without weight decay the active subspace freezes and generalization fails. This framework may explain why delayed generalization occurs, and why weight decay is strongly associated with grokking.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.