RAG Control Is Not Answer Fidelity: An Auditable Protocol for the Sufficiency–Fidelity Gap in Multi-Hop RAG
Abstract
Retrieval-augmented generation (RAG) fails on multi-hop questions when later support is indexed under an intermediate answer rather than the original query, and inference-time controllers are typically tuned to retrieve more, plan better, and stop smarter. We audit whether those control objectives actually track answer quality. Five-seed paired experiments with Holm correction are run on the official HotpotQA distractor split through PR-RAG, a training-free bounded protocol that exposes dependency-conditioned planning, evidence accumulation, a sufficiency gate, typed failure recovery, and per-call cost as logged, inspectable state. Three findings define what inference-time RAG control does and does not buy. (i) The headline: at the deployed operating point, 208/500 stopping decisions had evidence sufficient for the gold chain yet wrong answers, and 0/6 registered EM contrasts among the evaluated control variants survive Holm correction—evidence sufficiency is necessary but not sufficient for correctness. (ii) Control primitives act locally: planning improves chain coverage; verification changes stopping and cost, not answer quality; added retrieval budget raises coverage without EM, so evidence coverage and answer accuracy are separable control targets. (iii) The audit instrument is itself measurable: a blind two-annotator audit puts the offline gold-chain proxy's false-negative rate at 43.2%, while audit telemetry still contains a learnable sufficiency signal (question-disjoint offline calibrator). The protocol and per-question traces are released as a reusable auditing resource, and the structure replicates across five Llama-3.1-8B seeds and a dense retriever. Closed-pool scope, non-established EM superiority, and cross-dataset generalization remain open.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.