acceptodds
Under review as a conference paper at ICLR 2027

Beyond Relevance-Centric Retrieval: Rubrics-Oriented Document Set Selection and Ranking

Abstract

In RAG and deep research, answer quality depends on the entire evidence set a model obtains, and a top-k of individually relevant documents does not necessarily constitute a set that satisfies a complex information need. Yet evaluation has unfolded along the single axis of relevance, with discounted aggregates such as nDCG collapsing set quality into a sum of per-document relevance. Redundancy, conflict, and complementarity, properties that surface only at the level of the set, therefore never enter the measurement, though they are precisely what determines whether one set is better than another. We introduce SetwiseEvalKit, a document set evaluation benchmark organized as nine rubric dimensions across three levels, spanning both short-form and long-form scenarios and built from approximately 28K query-specific rubrics. A systematic evaluation of 12 rerankers yields three diagnoses: the strongest method covers fewer than 46% of the rubrics; the cross-document coordination dimensions are uniformly the weakest; and no single method leads in both scenarios. These diagnoses characterize what current methods lack, yet leave open whether the nine dimensions can conversely guide selection. We probe this with Rubric4Setwise, an exploratory configuration that supplies query-specific rubrics as selection criteria to Qwen3-8B. Without any fine-tuning, the resulting sets outperform many purpose-trained rerankers on downstream generation, draw on fewer documents and search rounds, and rank among the strongest in both scenarios. This corroborates the validity of the nine-dimension design and suggests that rubrics can pass from criteria for judgment into signals for selection, a direction that merits further study.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.