acceptodds
Under review as a conference paper at ICLR 2027

SkillNEO: Non-Elitist Evolutionary Optimization of Agent Skills

Abstract

Agent skills enable LLM agents to extend their capabilities without updating model weights. Skill evolution builds on this flexibility through continual skill refinement. Current methods use execution feedback to propose revisions, then select versions for further development based on current performance. However, a partially repaired skill may remain low-scoring until later edits resolve the remaining failures. To address this limitation, we introduce SkillNEO (Non-Elitist Evolutionary Optimization), which develops intermediate versions using comparative execution evidence. Its ordinary update selects parents from the current generation’s evaluated children without requiring parent-relative gains or a place among the best historical versions. Comparisons of repeated, parent-child, and sibling executions identify unresolved failures and useful procedures. Meanwhile, archived versions can be revisited with later evidence, and their traces remain available for subsequent edits. We evaluate task-conditioned skill development on 86 SkillsBench tasks across eight domains, using GPT-5.4 Mini in every model-based development role. Selected bundles raise Mini’s average reward from 49.1% to 66.4%, within GPT-5.5 and GPT-5.6 Luna’s original-skill performance range. On the 41 modified tasks, frozen-skill reuse raises Luna’s average reward by 10.5 percentage points; GPT-5.5 exceeds its original-skill reference by 10.6 points. Recorded searches show lower-scoring versions leading to later deployed bundles and supplying procedures to other branches. A weaker executor can thus develop skills reusable by stronger executors on the same tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.