SkillIF: Do Agents Follow Their Skills?
Abstract
Agent skills specify what an agent should produce, how it should perform the work, and which limits it should respect. Evaluating their use therefore requires examining whether agents follow these requirements during execution. We introduce SkillIF, a benchmark for systematically studying adherence to skill specifications. SkillIF pairs 96 curated skills across ten domains with executable tasks and derives 2,194 assessable constraints from their specifications. A shared taxonomy enables comparisons of the same requirement types across different skills, while deterministic checks and LLM judgments assess each constraint against execution evidence. Together, these components reveal which requirements agents repeatedly fail to satisfy and how those failures appear in their actions and outputs. Across combinations of language models and execution harnesses, procedural requirements emerge as a shared weakness, while explicit prohibitions are followed more consistently than scope limits and stopping requirements. Substantial violations also remain in runs accepted by the task judge. Analysis of successful runs with low adherence distinguishes missing content and unexecuted steps from outputs that fulfill requirements only partly or in a different form. SkillIF provides a systematic account of how agents follow reusable specifications and identifies concrete targets for evaluating and improving skill adherence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.