EvidenceGate: Evidence-Guided Commit-or-Repair Decisions for First-Hop Answers in Multi-Hop Question Answering
Abstract
Reinforcement learning has recently shown promise in improving searchaugmented reasoning. However, existing methods for open-domain multi-hop question answering (QA) may prematurely commit to a first-hop answer based on relevant yet insufficient evidence, causing errors to propagate through subsequent queries. To address this issue, we introduce EvidenceGate, an evidence-guided Commit/Repair framework that formulates first-hop answer commitment as a decision between committing the current candidate and performing additional retrieval for repair. To support this decision, we use a controllable search simulator to generate evidence with varying levels of support and train an evidence verifier to assess whether retrieved evidence supports a candidate answer. The verifier resolves clear evidence states, while a Commit/Repair value model handles ambiguous verifier rejections by predicting the cost-adjusted repair advantage from rollout-level final QA rewards. Extensive experiments on five multi-hop QA benchmarks demonstrate that EvidenceGate achieves state-of-the-art overall performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.