acceptodds
Under review as a conference paper at ICLR 2027

Spectral Bias in the Evolving Neural Tangent Kernel at Finite Width

Abstract

This paper shows how spectral bias behaves in the finite sample size and finite network width setting, for regression on the unit circle using a shallow ReLU network. Finite samples and finite width are the setting of practical training, which theory mostly overlooks. Spectral bias is the tendency of gradient-based training to learn the low-frequency part of a target before its high-frequency part. In the infinite-width limit the Neural Tangent Kernel (NTK) under the uniform measure on the circle is a convolution, so its eigenfunctions are the Fourier modes, and low frequencies, corresponding to larger eigenvalues, are learned faster. In this specific case, the limiting kernel and its spectrum are known in closed form. For a network of finite width, the NTK is computed from its actual parameters, so it is random at initialization and changes during training. Using the limiting kernel as a reference, we can measure two things separately for a finite-width NTK: how close its eigenspaces are to the Fourier modes, through the cosine similarity between the two subspaces (eigenspace alignment), and how large its average eigenvalue is on each frequency (spectral strength). With the kernel held fixed at initialization, i.e. the lazy training regime, we prove that on the low frequencies its eigenvalues and eigenvectors stay close to the limiting ones, with a deviation of in the number of training points and in the width , and the measured deviations decay at these rates. We observe that at smaller widths the neural network with evolving weights reaches a lower training loss than the same network trained in the lazy training regime, and it is correlated with an increase in the spectral strength of the lowest frequencies while eigenspace alignment degrades at higher frequencies. The increase persists even for a target from which these low-frequency modes are absent, and we prove that the expected rate at which the finite-width NTK changes at initialization does not depend on the target, so we believe the strengthened directions are set by the network's input embedding and bias, architecture, and parameter distribution. A formal theory of the evolving-kernel regime remains open.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.