acceptodds
Under review as a conference paper at ICLR 2027

State Space Attention: Learning Context-Dependent Memory Retention

Abstract

The key–value cache used by standard attention grows linearly with context length. Bounding this cache raises a central question: what should a sequence model retain, and how can it learn to make those decisions to support long-context reasoning? We introduce State Space Attention (SSA), an attention architecture that learns to maintain a bounded set of past representations and reads them using standard attention. A learned scorer repeatedly rescores all chunks in light of newly arriving context and the other contents of memory. We train the scorer and reader jointly, gradually moving from fully differentiable, softly weighting candidate memories to retaining a hard, fixed number. Moreover, the scorer receives training feedback based on whether retaining otherwise discarded information would improve the model’s predictions. Across our evaluated retrieval and reasoning benchmarks, SSA maintains full-attention Transformer-level performance while its active inference memory remains independent of sequence length. We compare SSA with recurrent models, bounded-memory attention methods, and full-attention Transformers across retrieval, variable tracking, multi-hop reasoning, and language modeling. We further pretrain 400M-parameter models on 15B tokens, and measure the associated computational and memory costs along with language modeling performance. Together, these experiments examine the capabilities and limitations of learned memory retention for long-context sequence modeling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.