Retrieval Layers: Fast End-to-End Retrieval with Sparse Attention
Abstract
Language model applications commonly require handling massive amounts of external information, for example querying internal documentation, reasoning over codebases, or user personalization based on past interaction history. Retrieval-augmented generation (RAG) is the dominant mechanism for grounding models in such external information. However, standard RAG frameworks are post-hoc scaffolds around the model, using a brittle lookup that cannot be trained end-to-end. Furthermore, RAG incurs significant computational overhead from prefilling large volumes of text, in turn slowing down subsequent generation by increasing the context length. To address these limitations, we introduce Retrieval Layers, end-to-end learnable sparse cross-attention layers that conduct fast lookup of external text in latent space, rather than prefilling retrieved text into the model’s context. Retrieval layers can be integrated into existing language models with minor post-training cost. We insert Retrieval Layers into an off-the-shelf Qwen3-4B and finetune it with custom SFT and RLVR datasets. On challenging multihop QA benchmarks, Retrieval Layers achieve Pareto-dominance on an accuracy versus end-to-end latency tradeoff in comparison to RAG baselines. We also show stronger performance than the Pareto frontier on long conversation reasoning and demonstrate that Retrieval Layers can extend the needle-in-a-haystack ability to 30M tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.