acceptodds
Under review as a conference paper at ICLR 2027

Skill: Fully Agentic Co-Evolution of Policy and Skill Library

Abstract

Skills have become the native carrier of procedural knowledge in deployed agents. However, most skill learning is misaligned with deployment in three respects, the artifact it produces, the pathway through which skills are invoked, and the tasks on which it is validated. The few fully agentic efforts evolve only one side, training the policy against a fixed skill library or evolving the skill library under a frozen model. Yet a skill matures on neither side alone. We present Skill², a fully agentic co-evolution framework in which a policy and its skill library advance together by alternating two separately trained roles around one skill library. The actor learns whether, when, and which skill to invoke through the native harness with Fork-GRPO, which forks a control branch at each skill decision so that credit lands on the tokens that made it. The extractor learns to write and maintain skills with counterfactual-utility policy optimization, which rewards each edit by the paired gain the current actor obtains from the revised skill library with dual-level credit assignment. Extensive experiments on seven real long-horizon agentic benchmarks demonstrate that Skill² is (1) effective, surpassing baselines with relative gains of up to 46.7% on terminal tasks and 59.8% on web tasks, (2) scale-consistent, with its average gain growing from Qwen3.5-4B to Qwen3.5-9B, and (3) transferable, lifting four larger frozen models from three families by up to 13.0% through the evolved skill library alone. Our code is available here.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.