acceptodds
Under review as a conference paper at ICLR 2027

CAN RETRIEVAL FAILURES DRIVE GRAPH EVOLUTION? SELF-IMPROVING GRAPHS FOR EVIDENCE RECOVERY IN MULTI-HOP QA

Abstract

Frontier models are increasingly expected to be self-improving, refining their reasoning, tools, and environments rather than remaining static after training. However, optimization typically targets only the policy, while the environment remains fixed. Multi-hop question answering makes this gap concrete: answering requires retrieving, connecting, and verifying evidence across documents, yet existing retrieval-augmented and graph-based systems operate over fixed retrieval environments, and reinforcement learning (RL) approaches typically rely on terminal answer rewards. Such rewards provide limited guidance on which retrieval decisions contribute to success or failure, and policy optimization alone cannot repair deficiencies in how the environment represents and connects evidence. We present SIG-R1, a self-improving graph method that combines retrieval-step rewards with failure-driven adaptation of a searchable hypergraph. SIG-R1 assigns process-aware credit to individual retrieval steps, diagnoses recurring failures, and admits only corpus-supported edits. Experiments show that, with Qwen3.5-9B, SIG-R1 improves macro-F1 by 6.4 points over Graph-R1, including a 12.0-point gain on HotpotQA. Three-seed ablations show higher mean F1 with graph evolution and dataset-dependent additional gains from process rewards. These results support coupling retrieval-policy learning with evidence-constrained environment adaptation. Our code is available at https://anonymous.4open.science/status/SIG-R1-DC07.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.