Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
Abstract
Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We formulate attention recall as a spherical-packing problem and introduce Kernelized Linear Attention Activations (KATA), a framework whose feature maps are derived from first principles certifying attention activations on self-dual homogeneous cone. Building on this observation, we show that rank-one positive semi-definite (PSD) features offer a favorable capacity–interference tradeoff. KATA also recovers a parameter-free convex output gate and characterizes associative capacity through the Welch interference floor. We implement KATA as fused Triton kernels at two operating points: a flash-attention-style forward up to FlashAttention-2 throughput, and an exact chunked-state form that reaches FlashAttention-2 forward throughput at k tokens. On long-range MQAR and repeated-key overwrite, several KATA variants outperform Gated DeltaNet, with parameter counts and state sizes reported alongside accuracy. Induction preserves near-perfect recall, while kernel benchmarks show that the maps can be implemented efficiently. KATA retains MQAR at a out-of-distribution length, approaching the softmax with roughly one quarter of the KV-cache entries. Experiments on 340M-parameter LLMs reveal a feature-dependent fluency trade-off and clarify how positional embeddings, delta rules, and decay gates interact with feature geometry.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.