Repetition or Corroboration? Evidence Selection for Retrieval-Augmented Generation
Abstract
Retrieval-augmented generation (RAG) ranks passages by relevance, but agreement among passages does not guarantee independent evidential support. Several passages may repeat one evidence basis, causing redundant support and discarding useful corroboration from distinct sources. We introduce BasisDPP-RAG, which separates the query-conditioned claim from its evidence basis and combines them in a factorized, quality-aware determinantal point process (DPP) for budgeted evidence selection. The resulting objective penalizes same-basis repetition while preserving agreement across distinct bases. Experiments on AVeriTeC, NLI4CT, and SciFact yield 84.84% average Exact Match, 11.17 percentage points above the strongest baseline, while reducing NLI4CT basis redundancy from 57.76% to 7.89% and improving robustness to adversarial distractors. These results show that reliable RAG should measure evidential support by distinct bases rather than passage count.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.