Coverage Is Not Utility: An Exact Ceiling on Selective Retrieval
Abstract
Retrieval-augmented generation degrades when the corpus does not cover the query, and the usual remedy is a gate that estimates retrieval confidence and falls back to parametric generation. Coverage is a useful intermediate variable, but it does not determine whether retrieval helps. Using HotpotQA-distractor and MuSiQue, which ship gold and distractor paragraphs for every question, we tog- gle evidence availability while holding question, distractors and pipeline fixed, obtaining ground truth for both coverage (is the evidence indexed?) and utility (does grounding beat answering from parameters?). The two diverge: grounding helps on only 29.4% of questions whose evidence is indexed. Because we con- trol coverage, we can evaluate a gate that knows it exactly, and it reaches only AUC 0.698 at predicting utility on HotpotQA and 0.757 on MuSiQue. This ceil- ing binds every score that carries no information about utility beyond coverage, cannot be raised by recalibration, and decomposes in closed form into two rates measurable without building any gate: how often grounding helps when evidence is and is not indexed. Reranker confidence, the standard signal, detects coverage (AUC 0.734) but predicts utility at only 0.602–0.675 and carries no detectable utility signal within the covered set, so the ceiling binds it. Asking the genera- tor whether the passages suffice does carry such signal on HotpotQA and reaches 0.639–0.724, though its advantage over the reranker is significant on one of three backbones and absent on MuSiQue; we offer it as an existence proof, not a better gate.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.