HOLA: Hadamard Higher-Order Linear Attention
Abstract
The attention mechanism is an important reason for the success of transformers. It relies on computing pairwise relations between a query and all key tokens and applying softmax normalization. The expressiveness of softmax stems from the exponential kernel, which, however, causes quadratic computational complexity. Linear attention replaces the exponential by a single inner product of feature maps. In terms of the Taylor-series expansion of the exponential function, this amounts to a first-order approximation. Prior work has shown that the second-order term can also be computed in linear time. In this paper, we introduce Hadamard Higher-Order Linear Attention (HOLA), which extends this construction in three directions. It supports efficient factorizations of arbitrary-order terms rather than only the second-order term. The shared symmetric feature map used by prior works is replaced with distinct maps for each factor in the higher-order Hadamard product. This relaxes a low-rank constraint on the state matrix. HOLA amounts to standard linear attention at order one, includes the prior second-order construction as an exact special case, and is compatible with existing mechanisms such as gating, decay, and delta-rule updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.