acceptodds
Under review as a conference paper at ICLR 2027

SkillAlign: Task-Aware Skill Evaluation through Capability–Quality Alignment

Abstract

Agent skills provide a lightweight mechanism for improving agent behavior through reusable procedural knowledge, without modifying model parameters. Yet as skill libraries grow, a fundamental question remains underexplored: what makes a skill useful for a particular task? Existing evaluation typically treats skill quality as either an intrinsic property of the skill or an empirical outcome measured through downstream execution. The former overlooks heterogeneous task requirements, while the latter is costly and offers limited guidance for skill selection and improvement before execution. We argue that skill utility should instead be characterized by the alignment between what a task requires and what a skill provides. We introduce SkillAlign, a task-aware framework for modeling this alignment prior to execution. SkillAlign represents tasks through six functional capability requirements and skills through six structured quality dimensions, connected by a capability-to-quality mapping. It computes capability-level task–skill matches using task-specific importance and primary-axis gating, preventing secondary qualities from compensating for deficiencies in capability-critical dimensions. The resulting alignment score supports skill ranking and selection, while capability-level deficiencies provide diagnostic signals for skill optimization. We evaluate SkillAlign across WebShop, ALFWorld, SWE-bench, and WebArena using complementary evidence from task profiling, controlled skill ranking, pairwise selection, quality-guided optimization, cross-model downstream execution, and mapping robustness. The results show descriptive variation among the assessed task profiles, task-conditioned evaluation discriminates high-quality Skills associated with successful executions from low-quality Skills associated with failed executions more effectively than holistic judgment or task-agnostic quality aggregation, and optimized skills consistently improve downstream performance across the tested agent models and benchmarks. Together, these findings establish task–skill alignment as a practical basis for pre-execution skill evaluation, selection, diagnosis, and optimization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.