acceptodds
Under review as a conference paper at ICLR 2027

Beyond State Capacity: Diagnosing and Improving Multi-Slot Linear Attention through Selectivity

Abstract

Multi-slot linear attention is commonly understood as expanding recurrent-state capacity, but this view alone cannot explain why routing designs with comparable state capacity differ in retrieval performance. We show that multi-slot routing is not merely state-capacity expansion: it factorizes a query-conditioned selection operator over historical tokens. Under this view, routing design determines the operator's usable expressivity even at fixed state capacity. We identify two structural restrictions in existing methods: cross-token selectivity leakage, where selecting one token also emphasizes others sharing its slot, and tied selectivity, where coupled read and write routing limits which selection patterns can be expressed. These diagnoses lead to SimpleRoute, a simple routing mechanism that removes both restrictions to improve selective retrieval while retaining linear complexity in sequence length. Trained from scratch with 1B parameters on 50B tokens, SimpleRoute achieves the highest average commonsense and real-world retrieval scores among the compared large-scale recurrent models, including 1.3B models trained on twice as many tokens, and improves real-world retrieval by 6.35 average points over the strongest baseline at the same scale and token budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.