acceptodds
Under review as a conference paper at ICLR 2027

Scoring a Set, Not Summing Passage Scores: Effective and Efficient Set Retrieval

Abstract

Multi-hop question answering requires retrieving multiple evidence passages whose usefulness often depends on one another. Conventional retrievers either rank passages independently or construct evidence sequentially through locally supervised next-passage decisions. Sequential conditioning captures some cross-passage dependencies, but its local extension scores do not provide a common criterion for comparing complete evidence sets of different compositions and sizes. A natural alternative is to rank evidence sets directly, which requires an explicit scorer over arbitrary query-set pairs. We therefore formulate multi-hop retrieval as direct ranking of candidate evidence sets through a learned query-set compatibility score . Following an energy-based structured prediction perspective, we learn this score by contrasting complete evidence sets against incomplete and noisy alternatives in the combinatorial set space. Once a gold evidence set is available, such alternatives can be generated automatically by adding, removing, or replacing passages, providing rich set-level supervision without additional annotation. To make inference over this combinatorial space practical, we introduce ParaSet, a lightweight scorer over precomputed passage representations for efficient set exploration, and SetCE, an expressive cross-encoder for set reranking. Across three multi-hop QA benchmarks and two first-stage retrievers, set-level retrieval consistently provides a complementary signal to passage-level relevance, becoming relatively more effective when less of the required evidence is directly connected to the query. Motivated by this complementarity, combining the two signals further improves downstream QA performance, outperforming both deeper passage-level retrieval and an ensemble of distinct passage-level retrievers. Finally, without invoking an LLM during retrieval, our method remains competitive with LLM-assisted evidence-selection pipelines while achieving up to faster evidence selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.