SkillDiver: Constraint-Guided Generation and Adaptive Differential Testing to Evaluate Skill Selection
Abstract
Large language model (LLM) agents increasingly use reusable skills, but selecting relevant skills becomes less reliable as skill libraries grow. Evaluating skill selection at scale remains challenging for two reasons: extending labeled benchmarks to new skills and usage scenarios requires costly task construction and label validation, while a task may admit multiple valid skill selections, making per-task correctness labels difficult to define. We develop SKILLDIVER, a framework that combines constraint-guided task generation with adaptive differential testing to evaluate skill selection without requiring per-task correctness labels. Constraint-guided generation extracts and samples fine-grained constraints from skill documents to generate tasks. Adaptive differential testing treats the selection from the full skill library as the selection under test and assesses it by comparing with selections from smaller reference pools. We use two types of reference pools: the target pool, which contains the skills used to generate the task, and the adaptive pool, which additionally includes skills in the selection under test. When the selection under test is not reproduced in either reference execution, SKILLDIVER reports a selection deviation as a potential selection error. Across two agents and four settings, SKILLDIVER exposes more selection deviations than GoS-Random-Direct, increasing the selection deviation rate by 9.8% to 90.4%. For validation on four labeled benchmarks, the differential module achieves an average F1 of 84.58% in identifying selections that disagree with benchmark labels, outperforming an LLM judge by 28.46 percentage points. With skill chains held fixed, constraint-guided generation increases within-chain semantic diversity by 76.6% on SkillsBench and 61.6% on R3-Skill. Overall, SKILLDIVER explores how task generation and differential testing can support automated skill selection testing without requiring per-task correctness labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.