SkillZip: Experience-Gated Macro Actions for Tool-Using Agents
Abstract
Language agents routinely rediscover the same multi-step procedures: search a repository, open the relevant implementation and tests, then format, test, and inspect a patch. Existing experience memories preserve such procedures as text that the model must interpret again. We introduce SkillZip, a training-free compiler that uses successful-trace support to select parameterized phase templates and admits each executable action only after sandbox replay establishes its observable contract. We evaluate the resulting action interface on 20 temporally held-out software tasks from 10 public repositories with two agent models and four matched conditions. Across 40 model–task pairs, SkillZip matches observed no-memory completion at 24/40 (paired 95% CI: to points) while reducing model-mediated tool decisions by 23.1% (repository-clustered 95% CI: 14.1–31.2% fewer; exact sign-flip ). It removes 31.2% of decisions versus an equally sourced textual workflow and 27.6% versus retrieved action sequences; against retrieval it also reduces API cost by 10.9% and wall time by 29.1%. Fresh regeneration reduces decisions by 29.6%, model cost by 18.4%, and raises success from 14/20 to 16/20. Across three paired Luna generations, the decision reduction is 35.4% (two-way repository–generation 95% CI: 19.6–45.0% fewer). Successful experience can therefore become a smaller interface for future decisions, not another block of context.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.