acceptodds
Under review as a conference paper at ICLR 2027

RSA: Recursive Sparse Attention with Hierarchical Deep–Shallow Memory and Sparse Activation

Abstract

Due to their linear complexity, linear sequence models excel in long-context modeling but underperform on tasks requiring long-range reasoning and retrieval compared to standard Transformers. Recent efforts using gating, delta rules, or multi-state models still mainly focus on short-range information, limiting effectiveness in complex long-sequence tasks due to insufficient information utilization and limited effective memory capacity. To address this, we propose a bio-inspired shallow-deep memory architecture, where multiple memory states are structurally interconnected and stacked to achieve exponentially large equivalent capacity: shallow memory stores coarse-grained representations and deep memory stores residual information. A correlation-guided sparse readout mechanism enables precise information retrieval from the stored states. Co-designing memory storage and retrieval, we introduce Recursive Sparse Attention (RSA), an innovative attention mechanism that bridges the gap between linear models and standard attention. Experiments show that RSA, to some extent, counters the skepticism that linear models may not scale to large parameter sizes, achieving performance comparable to or surpassing state-of-the-art linear and Transformer models on multiple benchmarks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.