acceptodds
Under review as a conference paper at ICLR 2027

SkillPrism: Taxonomy-Guided Evaluation and Cross-Skill Optimization of Agent Skills

Abstract

Agent skills extend the domain-specific capabilities of LLM agents. As the open skill ecosystem grows in scale and heterogeneity, evaluating skill utility and understanding how it varies across skill characteristics remain challenging. Existing benchmarks typically focus on selected domains or task sets, limiting their coverage of the broader skill ecosystem. We introduce SkillPrism-Bench, a systematic benchmark grounded in a multidimensional taxonomy that characterizes skills by artifact form, verification support, execution dependency, and control flow. The taxonomy guides skill sampling and task construction, enabling structured analysis of skill utility. Experiments across multiple models show that providing skills improves the average pass rate by 13.9 percentage points relative to the no-skill baseline, with distinct utility patterns across skill categories. Building on this taxonomy, we further introduce SkillPrism-Opt, which distills transferable editing strategies from source-skill trajectories and uses the target skill's taxonomy attributes to guide a one-step rewrite, without access to its test tasks or execution feedback. In multi-model evaluation on held-out skills, SkillPrism-Opt improves average pass@1 by 3.3 percentage points over original skills and by 5.2 percentage points over general rewriting. Together, our results provide a taxonomy and empirical foundation for systematic skill evaluation and demonstrate that cross-skill experience transfer can improve unseen skills without target-specific execution feedback.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.