acceptodds
Under review as a conference paper at ICLR 2027

Reinforcing Agents with Collective Skills

Abstract

We introduce SKILL2ENV, a scalable and aligned data recipe for modern agentic RL. Public Agent Skills provide workflow guidance across diverse domains, linking to real-world artifacts and packaging domain expertise and quality criteria. We develop a pipeline to turn 3.4k web-crawled and filtered Agent Skills into 8k executable terminal environments with both programmatic tests and behavioral rubrics. Through extensive RL experiments, we observe promising improvements in general agentic abilities. With 150 steps of RL training using a FlashREINFORCE variant, Qwen3.8-27B gains 9.0 percentage points on Terminal-Bench 2.1, and 7.1 percentage points on S2EBENCH, our hand-verified private benchmark featuring real-world agent use cases. Further analysis indicates increased behavioral alignment with original Agent Skills on similar problems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.