Canonical Regularisation of Wide Feature-Learning Neural Networks
Abstract
Wide neural networks in the feature-learning regime drive modern deep learning, and yet they remain far less studied than their kernel-regime counterparts. We consider a critical yet under-explored difference between these two regimes: the regulariser and prior implied by gradient flow training. This *canonical regularisation* property is well-studied in kernel regime networks – of all the infinite global minima, gradient flow selects exactly the vanishing ridge solution – and underpins the celebrated NN-GP correspondence, precisely allowing the modelling of noise during training. However, we prove ridge regularisation biases gradient flow in feature-learning regime networks, even in the infinitesimal limit of vanishing regularisation. Over training, ridge distorts the inductive bias of the network, with a particular damage done to pretrained networks where the implicit prior is informative. We resolve this by deriving the canonical regulariser from the training dynamics alone: under a single conditioning assumption, the initialisations that train to the same network form a curved surface, the basin manifold, and the optimiser's own kinetic energy, measured over the motions that preserve the training outcome, uniquely identifies ridge in the kernel regime and generalises it to *geodesic ridge* in the feature-learning regime. As a practical contribution, we propose *arc ridge* as a minimax-robust, scalable surrogate to geodesic ridge, revealing a deep relationship between early stopping and canonical regularisation across learning regimes. Finally, we demonstrate the consequences of our theory empirically on both image processing and NLP transfer-learning problems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.