Learning to Select Evidence from Coalition Feedback
Abstract
Language models using retrieved passages or stored memories must select useful evidence under a context budget. A common utility measure, leave-one-out deletion, misses useful passages when another passage supplies the same fact. On an exact-copy stress test, its supporting-passage AUROC is 0.50, compared with 0.84 for a Banzhaf coalition teacher that measures contributions across different subsets. We use this coalition feedback to train CAMS, a 462,849-parameter setwise selector whose credit objective requires no supporting-passage labels and whose online selection makes no generative language-model calls. Across five pool settings, credit supervision yields answer F1 close to passage-label supervision in most evaluated budgets. On a 100-passage hard-negative pool at budget eight, the pure credit-trained selector improves answer F1 from 0.608 to 0.647 over dense retrieval. Its forward pass takes 1.2 ms on an H20, compared with 3.6 s for LLM reranking and 8.4 s for prompted set selection; measured embedding, feature, and forward components total 65 ms. Coalition feedback thus turns an offline measure of contextual usefulness into inexpensive evidence selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.