acceptodds
Under review as a conference paper at ICLR 2027

EVE-Agent: Evidence-Verifiable Self-Evolving Agents

Abstract

Self-evolving search agents learn from the questions and answers they generate themselves; however, existing approaches typically use question difficulty as the primary training signal, providing little incentive to produce evidence that actually supports an answer. We ask whether self-generated evidence can serve as useful feedback for learning to answer questions both correctly and with justification. We introduce EVE-Agent, a self-evolving search agent that generates a candidate evidence span along with each question and answer. The proposer is rewarded based on how much the span increases a fixed solver's agreement with the generated answer when the span is provided versus when it is withheld, with search disabled in both conditions. This model-relative utility signal requires no external answer labels and directly measures whether the generated evidence is useful to the solver. The same span is subsequently reused as an evidence target for solver training, turning evidence generation into a reusable source of supervision. Across seven question-answering benchmarks, EVE-Agent consistently outperforms Dr. Zero in joint answer-and-evidence performance under the same backbone and retrieval pipeline. It also improves both average answer accuracy and judged evidence support. These results show that self-generated evidence can provide an effective training signal for self-evolving search agents rather than serving merely as an explanation appended to the final answer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.