Towards a Theoretical Foundation for Tool-Augmented LLM Skill Libraries: The Price of Composition, Propagation, and Safety Under Evolution
Abstract
Skill libraries, natural-language procedures that condition a frozen large language model (LLM), carry tools, and read and write persistent artifacts, are now a de facto standard for building agents. Their reliability, however, still rests on folklore: we cannot yet say how much return an agent loses by composing skills, or how long a safety guarantee survives as the library is edited. We construct a single end-to-end object that answers both: a certificate for a composed agent that bounds the return it loses to composition, measured against the best its frozen model could achieve on a fixed plan, with every term one the designer can drive down or measure and which stays valid over an explicit editing horizon. Two modeling choices make the bound well-posed. First, gaps are measured against the realizable optimum, the best conditioning of the fixed model on the given plan, quarantining base-model capability and plan structure as invariants. Second, a two-layer model separates a within-invocation decision process from a cross-invocation calculus on artifacts. We then assemble the certificate term by term: an exact decomposition names where return is lost; a propagation theorem bounds how per-skill errors compound along the plan, letting each skill rupture (fail discontinuously) with bounded probability rather than assuming Lipschitz behavior, and yielding a convergent/linear/exponential depth trichotomy with matching lower bounds, a checkpoint rule, and black-box-measurable constants; size and learnability bounds control the errors at their source; a two-regime safety result supplies the horizon; and a capability-flow margin demotes tool-combination danger to higher order in the rupture rate. Five algorithms make the certificate constructive. Empirically, a black-box harness recovers the propagation constants on a 332-skill, six-database SKILL.md suite; across five models and six domains the resulting deploy/reject verdict discriminates on both axes; and on real end-to-end LLM pipelines over a nine-model sweep the measured loss stays within the certificate in every case.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.