Retrieval Gains That Do Not Reach the Reader: Why Recall Misleads in Both Directions
Abstract
Retrieval-augmented generation is often evaluated by retrieval quality (e.g., Recall@), implicitly assuming that better retrieval yields better answers. We test this assumption directly. We introduce SCOPE (Schema-Checked One-Pass Expansion), which extends SIRA's validated query expansion to long structured documents (financial filings and papers). In a single call, SCOPE proposes additional search terms and likely table locations for the answer, each proposal is validated against the document before one search is issued. On questions from UDA, SCOPE improves Recall@10 over SIRA by points, and by points when the answer is in a table. Its retrieval is statistically indistinguishable from a five-turn ReAct agent (, with interval ) while using over fewer search tokens, and a one-turn agent matches it more cheaply still. In filings, much of the gain comes from a single instruction: preserve every year and fiscal period, since the year is often not the answer but the column it appears in. Despite delivering correct evidence more often, SCOPE does not measurably change answer accuracy. The net effect is churn, and evidence arrives for questions but disappears for that were previously answered well (), while newly served questions are answered worse (). Aggregate recall hides these transitions and can mislead in the other direction. At the same time, a re-ranker that reduces SCOPE by 5.4 Recall@10 points causes no measurable accuracy loss. A five-search agent improves accuracy with no measurable recall gain by collecting more varied context. We therefore recommend reporting evidence transitions (which questions gain vs. lose evidence) and accuracy conditional on evidence, alongside Recall@. We include document-level intervals throughout and report our preregistered targets and negative results in full.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.