Two Context Regimes with Opposite Signs: Measuring the Retrieval Tax on Natural Questions
Abstract
Retrieved context can raise average question-answering accuracy while losing answers that a model gives without documents. We measure both directions on 468 Natural Questions examples with a retrieval tax and a rescue rate, each paired against a closed-book cell under the same prompt contract, the system instruction that states how documents may be used. Under exact dense retrieval over 2.68 million passages, five passages lose 14.0% of closed-book-correct answers and repair 49.1% of the incorrect ones. Inside the same kind of genuine rankings, replacing every answer-bearing passage by the next non-answer passage raises the tax from 0.133 to 0.593 for Qwen3-8B in preregistered within-question runs, and by 0.33 to 0.46 on four models and two contracts; blinded semantic labels that adjudicate each question’s reference once for all its answers give +0.313, with identification bounds above zero. Masking only the gold-answer spans in the unchanged passages produces about half of the containment-scored increase on two models and, under semantic labels, exceeds a matched non-answer mask on complete cases and on questions with an adequate reference; both edits keep their effect at a fourfold reasoning budget scored on the full answer. Uninformative placeholder context costs 43–62% of known answers, but Qwen and Mistral decline by citing the context while Llama-3.1-8B mostly answers “I don’t know”, so on identical outputs two refusal detectors rank Llama-3.1-8B lowest and highest in known-question over-refusal.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.