acceptodds
Under review as a conference paper at ICLR 2027

Scaling Law of Label-NTK Alignment and a Tighter Convergence Bound in the NTK Regime

Abstract

Neural tangent kernel (NTK) provides a central framework for understanding the training dynamics of wide neural networks. In this work, we empirically identify a scaling law that often holds on real-world datasets, governing the alignment of training labels and label residuals with NTK eigenspaces. Specifically, the squared eigen-components of labels and label residuals scale approximately as and , respectively, where is the corresponding NTK eigenvalue. We theoretically explain this phenomenon through the Lipschitz continuity of real-world label functions. Finally, leveraging this scaling law, we derive a tighter convergence bound for gradient descent in the NTK regime. Unlike existing bounds that depend on the smallest NTK eigenvalue, which is typically extremely small and leads to pessimistic convergence rates, our bound depends on the entire spectrum and closely matches empirical training curves.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.