acceptodds
Under review as a conference paper at ICLR 2027

How Much Can You Project? A Neural Tangent Kernel Theory of Intrinsic Dimension

Abstract

Neural networks with millions of parameters are routinely trained within a surprisingly small effective parameter space, a phenomenon formalized by Li et al. as "intrinsic dimension", which has remained without theoretical explanation since its discovery. We provide the first such explanation for asymptotically wide networks, showing that intrinsic dimension emerges from the interplay of the Neural Tangent Kernel (NTK) and the Johnson–Lindenstrauss lemma. Specifically, we ask whether the NTK regime survives when training is restricted to a randomly chosen low-dimensional subspace of the full parameter space. We prove that it does when the subspace dimension exceeds a threshold of order : a block-structured random Gaussian projection preserves the NTK entry-wise throughout training with high probability as network width grows. Crucially, this threshold is independent of the ambient parameter dimension , which can be arbitrarily large. We complement this result with matching lower bounds: the dependence is unavoidable for random Gaussian projections, while no linear map can beat the threshold for the hardest gradient configurations. Thus, we identify a critical dimension above which random subspace training preserves the NTK regime and below which the same mechanism fails. We also identify a distinct threshold at : below it, the projected kernel is rank-deficient and a component of the training error is permanently frozen, with the effect determined by its alignment with the surviving eigenspace. Experiments on standard benchmarks support our theoretical predictions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.