acceptodds
Under review as a conference paper at ICLR 2027

Learning to Select Evidence from Coalition Feedback

Abstract

Language models using retrieved passages or stored memories must select useful evidence under a context budget. A common utility measure, leave-one-out deletion, misses useful passages when another passage supplies the same fact. On an exact-copy stress test, its supporting-passage AUROC is 0.50, compared with 0.84 for a Banzhaf coalition teacher that measures contributions across different subsets. We use this coalition feedback to train CAMS, a 462,849-parameter setwise selector whose credit objective requires no supporting-passage labels and whose online selection makes no generative language-model calls. Across five pool settings, credit supervision yields answer F1 close to passage-label supervision in most evaluated budgets. On a 100-passage hard-negative pool at budget eight, the pure credit-trained selector improves answer F1 from 0.608 to 0.647 over dense retrieval. Its forward pass takes 1.2 ms on an H20, compared with 3.6 s for LLM reranking and 8.4 s for prompted set selection; measured embedding, feature, and forward components total 65 ms. Coalition feedback thus turns an offline measure of contextual usefulness into inexpensive evidence selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.