Small Executors Derail on Intermediates They Produced Themselves: Compiling Skills into Tools They Can Act On
Abstract
Agent skills package procedural knowledge as natural-language instructions that an LLM agent loads at inference time, and are expected to carry over to any model without retraining. We find that they largely fail to reach small open executors. Offered through the standard skill channel, skills are often left unopened; placed directly in context, they are read and still fail. Small executors derail not on understanding the procedure but on the intermediate results they must produce themselves to carry it out. In controlled interventions, once these results are messy, the next step fails for executors of every size, and in a pre-registered handoff, a larger executor that takes over a small executor's state fails more often than after its own turns. We therefore compile each skill into tools that return clean intermediate results, once per task family from the public skill package and a single development instance, while the executor still decides which tools to call and how to finish. On SkillLearnBench, compiled tools raise a 4B executor from 9 to 53 of 80 held-out instances, above a 27B executor given the skill in context, and the gains replicate on a second model family. Written out as text, the same suite does not help: what small executors lack is not the knowledge in a skill, but tools they can act on.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.