acceptodds
Under review as a conference paper at ICLR 2027

MonoRoute: One Symmetric Low-Rank Search Scales Attention Memory to 1B Tokens

Abstract

We introduce MonoRoute to combine the answer quality of attention-based retrieval with high question-answering throughput. Memory Sparse Attention (MSA) integrates document selection and multi-hop reasoning within a transformer, but repeatedly searches the corpus in 18 layers using 1024-dimensional routing keys. MonoRoute replaces these searches with a single routing layer per hop; later layers reuse its selection while attending to their own document key-value states. Projection symmetrization constructs a shared query and key orthogonal projection, enabling rank reduction to 128 dimensions without retraining and cutting routing-key storage eightfold. Together, these changes reduce leading-order similarity-scoring memory and compute by 144×, allowing us to construct a corpus index at only 4 bytes per token, while preserving iterative retrieval. Across five QA datasets, MonoRoute exceeds FlashRAG on all four reported answer-quality metrics and achieves 3.1-6.5× its throughput on HotpotQA and 2WikiMultiHopQA at batch size 16, including 3.1× at one billion tokens. These results support attention-based retrieval as an accurate, high-throughput alternative to retrieval-augmented generation (RAG).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.