acceptodds
Under review as a conference paper at ICLR 2027

Differential Attention for Context Selection

Abstract

Large language models with long context windows can answer questions over full-length documents, but using the entire document is costly, slow, and not always more accurate. This paper studies long-context evidence selection: given a long document, a query, and a strict token budget, the goal is to expose only the most useful evidence to a downstream generator. Standard retrieval-augmented generation (RAG) methods retrieve semantically similar chunks using strong embedding models, but they can miss evidence that is sparse, structurally distributed, or weakly matched to the query. Prompt-compression methods, such as LongLLMLingua and CPC, reduce context in another way, but often do not explicitly model where query-specific evidence appears. We propose Selective Attention-Guided Extraction (SAGE), a training-free, plug-and-play context reduction framework. SAGE uses a lightweight frozen LLM during prefilling, without decoding or fine-tuning, to turn query-to-document attention into a document-level relevance heatmap. It then applies differential attention with a contrast query to suppress query-irrelevant background signals and isolate question-specific evidence. Finally, SAGE selects coherent high-scoring spans under the token budget and passes only the reduced context to a downstream generation LLM. Across QuALITY-hard, NarrativeQA, QASPER, and AIT-QA, SAGE improves the accuracy–token trade-off over strong embedding-based RAG baselines using Qwen3-Embedding and ColBERT, as well as prompt-compression baselines such as LongLLMLingua and CPC. On QuALITY-hard, SAGE achieves a top-4 leaderboard result using only a 30% context budget, showing that carefully selected attention-guided evidence can outperform much larger retrieved or compressed contexts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.