acceptodds
Under review as a conference paper at ICLR 2027

When Skills Fail to Compose: Benchmarking and Compiling Agent Skills

Abstract

Agent skills provide reusable procedures and resources, but supplying relevant skills does not ensure that an agent can coordinate them to complete a shared task. We introduce SkillsBench-Compose, a benchmark of 20 human-authored tasks that evaluates component requirements and cross-component relationships separately. This separation exposes failures in integrating otherwise successful components. We also propose SkillCompiler, which jointly compiles the supplied skills into one task-specific package before execution. It generates a coordinated workflow grounded in critical source rules, records guidance and resource choices in a structured intermediate representation, and packages the guidance with selected original resources. Across three language models, SkillCompiler improves on the highest-scoring baseline in each setting by 2.9–4.8 reward points on SkillsBench and 3.3–8.3 percentage points in full completion on SkillsBench-Compose, while using fewer total model tokens, including preparation. Ablations show gains over task-only guidance and independently rewritten skills, including when the rewritten skills are consolidated into a single package.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.