Shapley-Guided Set-Wise Reranking for Retrieval-Augmented Generation
Abstract
Retrieval-augmented generation (RAG) relies on reranking to select a small set of documents for a downstream LLM under a limited context budget. Existing rerankers either score documents individually or rank candidates jointly, but lack a principled way to model both the utility of a document subset and each document's contribution to that utility. We formulate RAG reranking as a cooperative game, where documents are players and a coalition's payoff measures its usefulness for the query. We introduce a set-wise valuer to model this coalition utility and show that a local desirability condition on the valuer is sufficient to induce the desired document ordering under a symmetric semivalue. This result motivates a set-contrastive training objective that avoids explicitly computing attribution values during training. At inference, we instantiate the same local-comparison principle with a practical restricted regression attribution rule. Specifically, we sample small coalitions according to a truncated KernelSHAP-inspired distribution and fit an additive regression to their utilities. We further train an amortized estimator to predict these attribution scores in a single forward pass. Across five multi-hop QA benchmarks, our method outperforms strong pointwise, listwise, and setwise baselines in reranking recall and downstream answer quality in most cases, while the amortized estimator retains most of the gain at the cost of a standard pointwise reranker, yielding a speedup over sampling-based attribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.