acceptodds
Under review as a conference paper at ICLR 2027

What Retrieval Buys Depends on Who Reads: Splitting Multi-Hop RAG Accuracy into Assembly and Reading

Abstract

Multi-hop questions name their endpoints but not the passages that connect them. Retrieval-augmented systems reach those passages either by building an entity graph offline or by retrieving iteratively with a model in the loop, and the two are compared by end-to-end accuracy. Accuracy, however, cannot tell whether a system failed to assemble the evidence or its reader failed to use it. We separate the two steps with two controls: at a matched passage count only the retrieval varies, and with identical passages only the reader varies. Whether the evidence holds the whole chain is scored apart from the answer. Together the controls split each reader's accuracy gap to gold evidence into an assembly share and a reading share. Measuring the steps apart changes verdicts. On MuSiQue, end-to-end accuracy ranks a single read of a graph index at least as high as an iterative loop. Yet at the same passage count the loop completes far more chains, and its passages raise accuracy once read after the search. The split also shows that the bottleneck differs by reader. On MuSiQue's deep questions most of the 8B readers' gap is reading: the index's top-ranked wrong passages cost open readers up to 32B more than the frontier model, at least twice as much by point estimate. Training moves one share and not the other: a small model distilled call by call from the loop's search learns to search, with no detectable reduction in its reading share. Running alone, it outscores the published systems we compare on the same MuSiQue questions. Retrieval gains are thus relative to the reader, and we propose evaluating them at a matched passage count and per reader.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.