CAN RETRIEVAL FAILURES DRIVE GRAPH EVOLUTION? SELF-IMPROVING GRAPHS FOR EVIDENCE RECOVERY IN MULTI-HOP QA
Abstract
Frontier models are increasingly expected to be self-improving, refining their reasoning, tools, and environments rather than remaining static after training. However, optimization typically targets only the policy, while the environment remains fixed. Multi-hop question answering makes this gap concrete: answering requires retrieving, connecting, and verifying evidence across documents, yet existing retrieval-augmented and graph-based systems operate over fixed retrieval environments, and reinforcement learning (RL) approaches typically rely on terminal answer rewards. Such rewards provide limited guidance on which retrieval decisions contribute to success or failure, and policy optimization alone cannot repair deficiencies in how the environment represents and connects evidence. We present SIG-R1, a self-improving graph method that combines retrieval-step rewards with failure-driven adaptation of a searchable hypergraph. SIG-R1 assigns process-aware credit to individual retrieval steps, diagnoses recurring failures, and admits only corpus-supported edits. Experiments show that, with Qwen3.5-9B, SIG-R1 improves macro-F1 by 6.4 points over Graph-R1, including a 12.0-point gain on HotpotQA. Three-seed ablations show higher mean F1 with graph evolution and dataset-dependent additional gains from process rewards. These results support coupling retrieval-policy learning with evidence-constrained environment adaptation. Our code is available at https://anonymous.4open.science/status/SIG-R1-DC07.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.