acceptodds
Under review as a conference paper at ICLR 2027

SkillNet-Gym: A Dynamic Benchmark for Compositional Skill Learning

Abstract

Large language model (LLM) agents increasingly rely on skills to solve complex, long-horizon tasks. However, existing benchmarks for agent skills are mostly static, manually curated snapshots, making them poorly suited to the continuously evolving nature of real-world skill ecosystems. To address this gap, we introduce SkillNet-Gym, a dynamic benchmark for evaluating compositional skill learning in LLM agents. Grounded in real-world sources, SkillNet-Gym automatically builds a continuously evolving heterogeneous SkillNet, and uses it to synthesize domain-diverse, compositionally complex tasks. These tasks enable unified evaluation for skill construction and skill composition. Evaluating eight frontier models with two widely used agent harnesses reveals that directly compressing documents into skills can cause information loss and agents remain weak at retrieving and orchestrating skills in the wild. These findings show that effective skill use is not a one-shot augmentation mechanism, but a compositional pipeline of interdependent competencies. SkillNet-Gym provides a dynamic testbed for developing, diagnosing, and improving evolving skill centric agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.