SkillEvolver: Skill Learning as a Meta-Skill
Abstract
Agent skills, folders of instructions and scripts that an agent loads at inference time, are today static artifacts: authored once, by human curation or by one-shot generation from the model's own knowledge, and never improved by use. Methods that do learn skills from execution traces learn them from many instances of a task distribution. We study online skill learning, in which a single new task arrives and only a handful of trials can be afforded before its skill must ship. We propose SkillEvolver, a meta-skill, itself loaded like any other skill, that drives any CLI-agent to author, deploy and refine a skill for that task; it learns the skill's prose and code rather than model weights, so the result drops into any agent without retraining. Two design choices make a handful of trials enough: each trial follows a distinct strategy the agent writes itself, so that a small batch explores different approaches rather than resampling one, and, when a first batch leaves failures, the distilled skill is refined from how fresh agents fail while using it. On SkillsBench, skills learned by SkillEvolver outperform the human-curated skills in overall resolution rate on all three models we evaluate, raising the resolution rate to 57.9% on DeepSeek-V4-Flash (13.3 points absolute improvement), 62.2% on GLM-5.2 (7.2 points absolute improvement) and 65.5% on Claude Opus 4.6 (10.8 points absolute improvement). Even when all four trials of a first batch fail, a second batch solves the training task 64.4% of the time after distillation, against 29.7% when drawn blind, and the loop spends 6.57 exploration rollouts per task on average.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.