acceptodds
Under review as a conference paper at ICLR 2027

Not All Skills Help: Measuring and Repairing Agent Skill Libraries

Abstract

LLM agents can improve without weight updates by accumulating natural-language skills, but a plausible instruction need not help the task on which it is used. We propose ASSAY, a framework that connects skill-library repair and task-specific selection through shared execution measurements. Randomized masking estimates each skill's marginal causal effect under sampled skill contexts across development tasks. The resulting task-indexed profiles guide conditional rewrites offline and are combined over nearby development tasks to select skills for each test instruction. Task similarity retrieves development evidence; execution-based scores determine which eligible skills to suppress. On GPT-5.1 / AppWorld *test_normal*, ASSAY achieves 77.4% completion versus 70.8% for ACE with the same templates. On the same restructured library, task-conditioned masking outperforms count-matched random and global-score masking by 7.4 and 4.2 percentage points, respectively. Independently regenerating attribution yields 76.8%. Across seven models and four providers, the complete pipeline improves over the respective baselines on both AppWorld splits and for five of seven models on -bench retail. GPT-5.1 improves from 49.9% with ACE to 66.4% on *test_challenge*. These results establish the practical value of connecting skill generation to empirical curation and task-specific use.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.