AlphaSkill: Agentic Reinforcement Learning with Self-evolving Skills
Abstract
Agent skills extract interaction experience into high-level, reusable knowledge that is essential for reliably solving complex tasks. Nevertheless, the acquisition of suitable skills remains challenging: manually designing them introduces a critical scalability bottleneck, while static skills may lead to a progressive misalignment between accumulated skills and the agent’s evolving capabilities. In this work, we propose **AlphaSkill**, a reinforcement learning framework that enables self-curated skills to co-evolve alongside the agent's improving capabilities. In this paradigm, a single policy network functions as two roles: (1) a *Reasoner* that interacts with the environment to solve tasks, and (2) a *Skill Curator* that maintains a skills library by creating or refining skills based on the frontier of the Reasoner’s capability. The Reasoner is rewarded for solving challenging tasks in environments, while the Skill Curator is optimized according to whether its curated skills contribute to the Reasoner's successes, thereby creating a virtuous loop of co-evolution between task solving and skills curation. Extensive experiments on various agentic environments and reasoning tasks demonstrate that AlphaSkill achieves consistent gains compared to its counterparts. Moreover, these self-curated skills can further improve test-time performance when used as in-prompt hints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.