Test What Matters: Learning to Verify Updates in Self-Evolving Agents
Abstract
Self-evolving language agents improve by updating instructions, memories, and skills, but identifying beneficial updates incurs substantial verification cost. When candidates are evaluated using fixed task sets or random sampling, the selected tasks may not match those actually affected by the update. This mismatch can waste evaluations on unaffected tasks and leave critical improvements or regressions undetected, leading to incorrect update decisions. We introduce BRACE, Bayesian Risk-Aware Candidate Evaluation, which selects validation tasks to inform candidate-specific deployment decisions. BRACE learns shared response structure offline from historical paired outcomes to guide validation task selection. During online verification, BRACE updates predictions from new paired observations and uses Decision Value of Information (DVI) to prioritize tests by their expected reduction in deployment risk. A calibrated posterior-risk threshold governs stopping. We evaluate BRACE in SkillOpt and EvoSkill on SpreadsheetBench-Verified and DocVQA across four language-model configurations. With Qwen3.5-4B, BRACE reduces fresh online validation episodes by 54.0–82.3% relative to validation with a determinantal point process (DPP) coreset while achieving higher mean terminal-run success rates. Cross-model comparisons also show lower total validation cost and higher mean terminal-run success rates than host-native validation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.