Weight Anisotropy in Mean-Field Theory: Learning on Isotropic Data
Abstract
with low-dimensional target structure (canonical examples include -sparse parity and sparse multi-index models), yet their neural tangent kernel limits require samples in the ambient dimension . We trace this gap to a single mechanism: during training, networks develop strong anisotropy between weight coordinates aligned with the task and those that are not. We call this input feature selection (IFS) and show, through analysis of stochastic gradient Langevin dynamics, that it arises from coordinate-dependent effective regularisation that kernels structurally cannot exhibit. Mean-field (MF) theory is the natural interpretable framework for feature learning beyond the kernel regime, but standard MF tracks only first moments of the weight distribution and so cannot represent IFS. We introduce MF-ARD, which augments MF with coordinate-wise precisions through automatic relevance determination. With this single additional set of order parameters, MF-ARD (i) captures the sharp generalisation transitions of SGLD-trained networks on -sparse parity and single-index models, and (ii) provably breaks the curse of dimensionality: its phase-transition threshold depends on the intrinsic task dimension rather than the ambient dimension .
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.