acceptodds
Under review as a conference paper at ICLR 2027

The Value of Evidence Selection Depends on Reasoning Effort

Abstract

Evidence selection reduces input length but also changes the information available for reasoning. Its accuracy and cost benefits may therefore depend on the reader’s reasoning effort. We test this dependence in multi-hop retrieval-augmented question answering with the model, prompt, retrieved passages, and selected contexts fixed. On held-out 2Wiki questions, a learned selector improves exact-match accuracy by 4.4 percentage points at minimal effort but reduces it by 3.0 points at medium effort. Despite the same 67.3% prompt-token reduction, reader-cost savings fall from 60.0% to 8.7%. To examine the role of evidence content, we restore omitted support in an annotation-guided intervention across three benchmarks. Relative to adding a similar-length non-support passage, restoration improves medium-effort accuracy by 9.3 points and reduces mean reported reasoning tokens by 42.2%. We extend this investigation to information computed within a source-only retrieval pipeline. Grounded execution reuses intermediate answers already extracted to guide search; exposing rather than withholding these answers, with acquired passages fixed, improves minimal-effort accuracy by 5.3 points on a three-benchmark development panel without another model call. Reuse also lowers medium-effort total API cost, including planning and extraction, while its accuracy effect is unresolved. These effort-dependent benefits motivate joint configuration selection. We evaluate joint calibration of existing evidence policies and reader effort on held-out questions.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.