Found but Forgotten: The Evidence Assembly Gap in Retrieval-Augmented Generation
Abstract
Building a strong retrieval-augmented answer generation system requires more than finding relevant passages: the system must also deliver the right combination of evidence to the reader. We began with a frozen HippoRAG2 core because it provides a strong graph-based retrieval foundation and exposes frontier traces that indicate where its search may be incomplete. Our goal was to determine how far this foundation could be pushed by improving the search-and-selection pipeline around it—and, in doing so, identify which bottlenecks remain once retrieval itself is already strong. We find that a candidate pool can contain all the evidence needed to answer a question, within the reader’s passage budget, yet the final context can still omit a crucial piece. We call this the evidence assembly gap. To study it, we separate three stages: whether sufficient evidence is available among the candidates, whether the complete evidence set is preserved during selection, and whether a bounded reader successfully uses it. A compositional-witness framework explains why passage-level recall does not guarantee recovery of a complete justification, and why simply giving the system more evidence need not improve its answers. We introduce CAST, a graph-assisted method built on top of a HippoRAG2 core with joint evidence selection. On 1,000 MuSiQue questions, the method improves token F1 by 11.89 points over the frozen native baseline and by 6.10 points over same-model pointwise selection, while keeping the reader fixed at five passages. Joint selection also reduces scenarios in which a feasible support set is available but discarded by 55.1%. Across seven packaged cohorts, whole-pipeline mean F1 improves on every dataset, with six nominal paired intervals above zero. Further analysis shows that preserving evidence and successfully answering are distinct problems, and that the value of additional reader context depends on how that context is constructed. Finally, model-assisted audit reveals different remaining bottlenecks across benchmarks: some exhibit near-complete annotated retrieval but substantial reference or scoring disagreement, whereas long-context QA continues to suffer from evidence-selection failures. Together, these results provide a concrete method for turning broad retrieval into compact, usable evidence and a stage-specific account of where further algorithmic and benchmark progress is needed due to dataset saturation and other issues, along with disaggregating different viable strategies to approach different types of retrieval problems in the future.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.