acceptodds
Under review as a conference paper at ICLR 2027

Learning to Compose Evidence with Counterfactual Con- tinuation Values

Abstract

Retrieval-augmented question answering selects passages that are consumed jointly, yet common selectors score passages independently or optimize hand-designed set proxies. We study counterfactual continuation-value learning for bounded evidence selection. Given a small candidate pool and a frozen reader, we evaluate every feasible passage subset, cache its an- swer utility, and derive action targets from the best utility reachable after adding each passage. A compact selector learns these targets by regression and composes a context sequentially, with an explicit stop action. On Hot- potQA, it improves over BGE Top-4 on two disjoint 64-question test sets across two training seeds. On a fresh 256-question sample, the budget-5 selector reaches 53.9 EM / 65.8 F1 versus 47.7 / 59.7 for BGE Top-4 and 50.8 / 62.7 for All-8. Without retraining, it transfers to two independent 2WikiMultiHopQA samples: 38.3 / 46.1 versus 31.6 / 40.3 on 256 ques- tions, and 34.9 / 41.8 versus 30.3 / 35.3 on a disjoint 768-question sample. On a matched 768-question 2Wiki pool, a GRPO refinement reaches 35.29 EM / 43.11 F1, exceeding E5 Top-5 by 2.21 / 4.59 points at similar token cost. Candidate pools retain annotated supports, so these results evaluate composition conditional on support availability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.