Kernel Velocity: A Local Certificate for Feature Learning beyond the NTK
Abstract
The initial neural tangent kernel (NTK) does not reveal whether a network will keep that kernel fixed or learn task-dependent features. We isolate the missing local quantity: the task-effective kernel velocity (\dot K_0 r_0), where (K_t) is the empirical NTK and (r_t) is the training residual. For any smooth network trained by preconditioned gradient flow, we prove that the nonlinear predictor and its frozen-NTK trajectory (f_t^K) first separate as [ f_t-f_t^K =-t^22N\dot K_0r_0+O(t^3). ] We then construct a centered two-layer scaling family whose members have exactly the same initial predictor and NTK. Their kernel velocities obey the exact finite-width law (\dot K_0^(\alpha)=m^\alpha-1\cV_m); a nonzero population limit yields a vanishing lazy velocity at (\alpha=1/2) and order-one velocity at (\alpha=1). A rank-one linear model gives an exact adaptive-kernel closure and a monotone law for oriented teacher alignment. Experiments recover the predicted width exponents, verify second- and third-order trajectory expansions in two-layer and deep networks, and identify task-aligned nonlinear features. At a common training time, the adaptive scaling lowers test mean squared error from (0.294) for the matched frozen NTK to (0.160), while its feature spike has (0.93) teacher alignment. A stopping-time stress test separates this acceleration from a predictive advantage: on the three harder teachers, adaptation also beats the minimum test error on a dense exact-flow grid. Kernel velocity is a computable local certificate of departure from fixed-kernel training, not by itself a global guarantee of useful representation learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.