Linearized Attention Cannot Enter the Kernel Regime at Any Practical Width
Abstract
Influence functions are now widely used to audit transformer training data, relying on the unverified assumption that attention reaches the neural tangent kernel (NTK) regime. Softmax attention admits no exact NTK, so this question is studied in this paper through linearized attention, the standard tractable proxy. This proxy is shown not to enter the kernel regime at any practical width. An exact correspondence is established between parameter-free linearized attention and a data-dependent Gram-induced kernel whose condition number is the cube of the input Gram's. With the effective condition number of the input Gram matrix, a width suffices for convergence, and a matching necessity bound makes the exclusion two-sided. For natural image datasets, with (MNIST) and (CIFAR-10), even the tight requirement exceeds , more than ten orders of magnitude beyond the parameters of the largest known architectures. Influence malleability is introduced to measure the consequence. Linearized attention exhibits 3–9 higher malleability than ReLU networks under adversarial data perturbation, with the gap set by dataset condition number and task. For trainable query, key, and value projections, an exact NTK block decomposition yields a precondition floor . Non-convergence then persists at exponent two for every dataset studied. The cubic mechanism itself is erased under standard independent initialization and survives only under an explicit, checkpoint-measurable condition. This precondition is measured above the non-convergence threshold across more than twenty pretrained language, vision, and audio models, the setting where influence methods are deployed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.