Skill-R1: Agent Skill Evolution via Reinforcement Learning
Abstract
Agentic large language models increasingly rely on skills, reusable natural-language procedures that guide planning, action, and tool use. Current practice improves skills through prompt engineering or by fine-tuning the task LLM, both of which are costly, model-specific, and often infeasible for closed-source models. Effective skill optimization is moreover recurrent rather than one-shot: a useful skill must improve rollouts under the current conditioning, and a useful revision must turn observed outcomes into a better skill for the next round. We propose Skill-R1, a reinforcement learning framework that addresses skill optimization as a bi-level problem coupling within-generation execution quality with across generation skill improvement. Skill-R1 trains a lightweight skill generator conditioned on the task context, prior rollouts, and verified outcomes to produce skills that steer a frozen task LLM, preserving compatibility with both open and closed-source models. The generator is optimized via a bi-level group-relative policy optimization objective: an intra-generation advantage compares skill–rollout pairs within each generation, and an inter-generation advantage rewards population-level improvement across generations. On agent benchmarks with verifiable rewards, Skill-R1 yields substantial gains over prompt evolution, and Best-of-N sampling methodologies, with the largest improvements on complex multi-step reasoning benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.