BRAID: Evolving Evaluators for Reliable and Effective Self-Evolving Agents
Abstract
Self-evolving LLM agents iteratively modify themselves from task experience, making reliable and actionable evaluation signals critical to efficient improvement. Existing systems, however, largely evolve the Task Agent while leaving the Evaluator that diagnoses failures and provides feedback fixed. To address this issue, we introduce Branch-Realized Adaptive Interventional Dual-Evolution (BRAID), a training-free framework that is agnostic to the underlying Task-Agent search topology and treats the Evaluator as an explicitly evolvable agent with an independent evolutionary lineage. BRAID periodically generates Evaluator challengers from accumulated task trajectories, modification histories, and ground-truth outcomes. A challenger replaces the incumbent only when it passes an Interventional Reliability Gate and, through Challenger Branch Realization, produces a Task Agent branch whose realized branch utility exceeds the incumbent utility baseline. This selection goes beyond stand-alone evaluation accuracy by measuring the downstream utility of the Task Agent modification induced by each challenger's feedback. We evaluate BRAID across multiple self-evolution topologies and tasks spanning executable code, repository-level software engineering, and mathematical grading. Compared with the corresponding underlying-system baselines, BRAID consistently improves held-out performance under matched search-action budgets, while Frozen-Evaluator controls isolate the additional gains from Evaluator evolution. These gains persist on Token–Performance curves after accounting for Evaluator-side computation, and ablations confirm the complementary roles of Evaluator evolution, the Interventional Reliability Gate, and Challenger Branch Realization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.