Sliced-Linear Attention: Beyond the Finite-Feature Bottleneck of Linear Attention
Abstract
To address the quadratic bottleneck of Softmax attention, many efficient attention mechanisms rely on finite-dimensional summaries of the context and, in the causal setting, on fixed-dimensional recurrent states. We introduce Sliced-Linear attention, which multiplies a standard Linear-attention kernel by a Laplace interaction between learned one-dimensional query and key coordinates. This removes the finite-rank constraint of finite-feature Linear attention while retaining exact quasi-linear computation through sorting and scans. We show that, for any prescribed finite key set, a single Sliced-Linear head can approximate arbitrarily sharp associative retrieval, whereas finite-feature Linear attention requires to maintain fixed relative error as memory size grows. In the causal setting, exact retrieval of arbitrary -dimensional values by a continuous recurrent model requires state dimension . Controlled MQAR experiments show the corresponding qualitative capacity separation. On Augmented LRA and in 124M-parameter causal language modeling, Sliced-Linear consistently improves over matched Linear attention, with larger gains after continuation to 16k and 32k context. With , it outperforms Linear attention with across all tested text settings. Measured forward-and-backward runtime scales near-linearly with sequence length.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.