acceptodds
Under review as a conference paper at ICLR 2027

SkillRein: Reusable Agent Skills Generation with Reinforcement Learning

Abstract

Recursive self-improvement calls for agents that improve not only task execution but also the mechanisms that drive future improvement. Motivated by this, we ask whether language agents can learn to turn their own execution experience into effective, reusable guidance. We introduce SkillRein, a reinforcement learning framework that trains a dedicated proposer to generate natural-language skills while keeping the task executor frozen. For each batch of related tasks, the proposer receives selected initial trajectories collected without skills and generates a shared skill set. The executor then reattempts the same tasks using these skills, with task outcomes and skill-use feedback jointly providing the reward for training the proposer. This directly optimizes the skill-generation policy for downstream execution performance. Across four search task categories, a trained 4B proposer raises overall success from 30.2% to 35.6% with a Qwen3-4B executor and from 18.6% to 30.1% with a Qwen3-8B executor. It also outperforms three general-purpose language-model proposers in overall success with both executors. These results show that reinforcement learning can improve an agent’s ability to generate useful guidance from execution experience, supporting learned skill generation as a building block for agent self-improvement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.