MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
Abstract
Long-term memory is fundamental to intelligence, yet enabling AI to process lifetime-scale information on the order of 100M tokens remains challenging. Full-attention architectures limit LLMs to roughly 1M tokens, and existing extensions such as linear attention, fixed-size memory, or external storage methods suffer from accuracy loss, growing latency, limited memory adaptability, or poor end-to-end optimization. These limitations restrict scalability and hinder applications including large‑corpus summarization, stable digital twins, and long‑history agent reasoning. We present Memory Sparse Attention (MSA), an end-to-end trainable, efficient, and massively scalable memory model framework. Through core innovations including scalable sparse attention architecture and document‑wise RoPE, MSA achieves linear complexity in both training and inference while maintaining exceptional precision stability, exhibiting less than 9% degradation when scaling from 16K to 100M tokens. Furthermore, KV cache compression, combined with Memory Parallel during inference, enables 100M tokens inference on GPUs. In addition, we propose a Memory Interleave mechanism that effectively facilitates complex multi‑hop reasoning across scattered memory segments. MSA significantly surpasses frontier language models, state-of-the-art (SOTA) RAG systems, and leading memory agents in long-context benchmarks. These results demonstrate that by decoupling memory capacity from reasoning, MSA provides a scalable foundation to endow general-purpose models with intrinsic, lifetime-scale memory. Model available at https://huggingface.co/Anoy123423123/MSA-4B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.