Test Time Computing for Evolving Agent Skill Optimization
Abstract
Large language model (LLM) agents increasingly rely on reusable skills that en- code procedural knowledge for specialized tasks. Recent work has begun to op- timize such skills automatically from execution feedback, but existing methods primarily focus on how to revise an individual skill, while the outer search over alternative refinement paths remains comparatively underexplored. Early revi- sion choices can therefore narrow subsequent exploration, and simply spending more optimization computation along the same trajectory may not recover alter- native directions. We introduce EvoBranch, a plug-in evolution framework that separates the native skill-revision mechanism from the search strategy around it. EvoBranch reuses an existing skill optimizer without modifying its revision pro- cedure, while broadening exploration through two complementary mechanisms. Branch-local optimization uses execution trajectories and branch-specific history to produce multiple targeted revisions from each parent skill, while global evolu- tion generates additional candidates across branches through mutation, crossover, and regeneration. Validation-based Top-M beam selection then determines which refinement paths receive further optimization. We evaluate EvoBranch with 3 skill optimizers, 3 LLM backbones, and 4 benchmarks. EvoBranch im- proves the average performance of all nine model-optimizer combinations, with gains of up to 51.4 percentage points .Compute-controlled comparisons further show that longer single trajectories and independent restarts do not reproduce these gains.Anonymous code is available at: https://anonymous.4open. science/r/EvoBranch-for-skill-optimization-DA3F/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.