acceptodds
Under review as a conference paper at ICLR 2027

Submodular Evidence Selection for Retrieval-Augmented Generation

Abstract

Retrieval-augmented generation (RAG) conditions a frozen generator on a -passage subset of a larger candidate pool. Relevance-based selection suffers two failure modes that a larger budget relieves only inefficiently: facet starvation, where the top ranks over-serve one semantic dimension of the query while its remaining facets scatter at lower ranks, and higher-order redundancy, where pairwise-dissimilar passages jointly span a low-dimensional subspace and crowd out multi-hop evidence—in both cases, extra budget buys redundant copies rather than the missing evidence. We propose Submodular Evidence Selection (SES), which maximizes a single monotone submodular objective combining relevance-gated coverage—a frozen language model decomposes the query into sub-queries, and direct relevance is fused multiplicatively into per-facet coverage credit, so coverage is earned only by passages that themselves answer the query—with geometric diversity, the log-determinant volume spanned by the selected embeddings, which penalizes redundancy of any order. Both normalizers are constants fixed before optimization, so greedy maximization inherits the classical guarantee exactly, with no surrogate step. Across four multi-hop QA benchmarks and three frozen generators, SES consistently improves exact match over relevance-based and diversity-aware baselines, composes cleanly with cross-encoder reranking, and adds only millisecond-scale CPU selection beyond one cacheable decomposition call per query.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.