Verification as a Budget: Posterior-Separation Scheduling for Self-Evolving Agent Skills
Abstract
Optimizing a natural-language skill document instead of model weights is a low-cost path for adapting frozen LLM-based agents to downstream tasks, and mainstream methods share one outer loop of target-model rollouts, optimizer proposals, validation and accept-or-reject decisions. This paper identifies a structural deficiency in that loop: validation is fixed as a full-scale step. Every candidate is evaluated once on the full validation set and the gate compares that noisy reading against the incumbent's recorded score, a score systematically inflated by selection bias whose raised threshold crowds out genuine late-training gains; rejected candidates cannot be resurrected and their measurements accumulate no evidence. Meanwhile target-model rollouts and optimizer calls, two operations with very different cost structures, are bound at a fixed ratio set by a step counter and insensitive to the current estimation quality. We propose VeriTree, which restructures the outer loop as a skill tree with uncertainty estimates. Each round performs either additional validation or the generation of a new candidate, and the switch is decided in closed loop by a two-arm criterion on posterior separation: validation stops and expansion proceeds once the leading candidate is statistically separated from its strongest competitor, or once separating them would cost more validation than the remaining budget allows. Validation targets are chosen by uncertainty-aware sampling, optimizer calls are subject to data gating, and delivery favors well-evidenced nodes; the whole behavior is controlled by a single parameter set from the task configuration. Across three optimizer–target configurations and five tasks, compared with SkillOpt and GEPA under identical accounting, VeriTree attains comparable or better accuracy, while reducing training cost to 10%–63% of SkillOpt's on every task and under every configuration. The results support the paper's central claim: validation in self-evolving agents should be treated not as a fixed full-scale step but as a scarce budget scheduled by posterior state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.