acceptodds
Under review as a conference paper at ICLR 2027

RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States

Abstract

Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor, a matrix, limiting the order of interactions that can be represented in the state. We introduce the RunningTensor, which generalizes this memory to an order- tensor, updated by a rank-1 outer product and read by contracting against vector queries. Order recovers linear attention; we study order as a proof of concept, retaining both recurrent and parallel forms while remaining linear in sequence length and improving working memory capacity from to . On synthetic multi-query associative recall, RunningTensor outperforms linear attention and SSM baselines. It also improves performance on time-series classification and forecasting, suggesting that higher-order recurrent state can provide useful additional memory capacity beyond matrix-valued state.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.