Skill: Skill Evolution as Policy Optimization
Abstract
Skills are increasingly generated and refined automatically, but the dominant sequential-refinement paradigm—execute the current skill, reflect on its rollout, and rewrite the whole skill—often degrades over rounds. We trace this failure to three missing ingredients from reinforcement learning: a baseline for the return, credit assignment to individual design choices, and a replay buffer that preserves evaluated skills and trajectories. These mechanisms are not optional heuristics: without a return baseline, dimension-level credit assignment, and a replay buffer, the effect of any design choice is unidentifiable from a single rollout. We therefore reframe skill evolution as optimizing the policy that writes skills, not the skill document itself. The actions are discrete design choices within a skill, and the policy parameter is a growing context of evaluated skills, trajectories, returns, and conclusions. We propose Skill, which performs controlled one-dimension-at-a-time comparisons: alternatives in a group differ only in the chosen dimension, so the rest of the group serves as a baseline; the return gap is credited to that dimension; and every evaluated skill remains available as a reusable control. The final skill is synthesized from choices that won such comparisons, while unsupported or null dimensions are removed. On SkillLearnBench, across three LLMs, Skill achieves 52.0% average accuracy, 9.4 points above the no-skill agent and 11.1 points above the strongest skill-generation baseline, and improves held-out transfer and judged skill completeness/determinism. Ablations confirm the contribution of controlled comparison, control reuse, and synthesis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.