Evidence in the Loop: Training Search Agents from Knowledge Graphs
Abstract
Training search agents increasingly relies on synthetic data to reduce human annotation costs. Most prior methods generate training data offline, leaving the task distribution fixed and unable to adapt to the agent's evolving capabilities. Self-evolving frameworks address this limitation by synthesizing data online as the agent improves. However, optimization remains largely outcome-driven, resulting in sparse supervision for intermediate search behavior.Recent work therefore incorporates sampled knowledge graph (KG) paths into self-evolving training, using relational structure to guide question generation and path entities as intermediate supervision. However, this entity-centric design can weaken both the effectiveness of training data and the supervision signal: generated questions may still be answerable even when parts of the factual chain behind it are not explicitly supported by the target text corpus, while entity-based rewards can conflate mentioning path entities with genuine search progress, even when no relevant evidence is actually retrieved to support the solution. To this end, we propose Evidence-in-the-Loop, a self-evolving framework that grounds both question generation and process supervision in retrievable corpus evidence. We sample multi-hop KG paths and retain only those whose triplets are supported by verified corpus evidence. The retained paths guide online question generation, while the associated evidence is reused to supervise the agent's search process. Specifically, we introduce two process rewards: evidence acquisition, which rewards retrieving verified supporting passages, and evidence utilization, which rewards using newly acquired evidence to guide subsequent search toward additional relevant evidence. Experiments on knowledge-intensive QA benchmarks show consistent improvements across models and initialization settings, with particularly strong gains on multi-hop QA. The trained Proposer also produces higher-quality questions, as reflected in both their validity and downstream training utility.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.