Reinforcing Agents with Collective Skills
Abstract
We introduce SKILL2ENV, a scalable and aligned data recipe for modern agentic RL. Public Agent Skills provide workflow guidance across diverse domains, linking to real-world artifacts and packaging domain expertise and quality criteria. We develop a pipeline to turn 3.4k web-crawled and filtered Agent Skills into 8k executable terminal environments with both programmatic tests and behavioral rubrics. Through extensive RL experiments, we observe promising improvements in general agentic abilities. With 150 steps of RL training using a FlashREINFORCE variant, Qwen3.8-27B gains 9.0 percentage points on Terminal-Bench 2.1, and 7.1 percentage points on S2EBENCH, our hand-verified private benchmark featuring real-world agent use cases. Further analysis indicates increased behavioral alignment with original Agent Skills on similar problems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.