SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
Abstract
Real-world tool-using agents operate over long-horizon workflows with recurring structure and diverse demands, where effective behavior requires not only invoking atomic tools but also abstracting and reusing higher-level tool compositions. However, existing benchmarks mainly measure instance-level success under static tool sets, offering limited insight into agents’ ability to acquire such reusable skills. We address this gap by introducing SkillCraft, a benchmark explicitly designed to stress-test agents' ability to form and reuse higher-level tool compositions, which we call Skills. SkillCraft features realistic, highly compositional tool-use scenarios with difficulty scaled along both quantitative and structural dimensions, designed to elicit skill abstraction and cross-task reuse. We further introduce a lightweight evaluation protocol that lets agents auto-compose atomic tools into executable Skills, cache and reuse them within and across tasks, thereby turning compositional skill acquisition into a directly observable, measurable behavior. Evaluating state-of-the-art agents on SkillCraft, we observe a clear capability gradient: stronger models exploit the same composition affordance to reduce token usage by up to 80%, while weaker ones fail to do so. Moreover, success rate positively correlates with skill execution ability, establishing compositional skill acquisition as a core, separately measurable capability of tool-using agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.