SkillCraft: Skill-Set Retrieval Beyond Ranking for LLM Agents
Abstract
Large language model (LLM) agents increasingly rely on external procedural skills to solve complex tasks, yet existing skill retrieval predominantly follows a relevance ranking formulation. We show that this formulation overlooks a critical gap: candidate discovery does not imply set recoverability. In multi-skill queries, target skills are frequently interleaved with irrelevant distractors, destroying rank-prefix recoverability. On SkillRet, while top-50 retrieval achieves 89.6% candidate completeness, 46.5% of complete queries remain source-rank inseparable, admitting a median of seven unavoidable intruders under contiguous prefixes. To resolve this bottleneck, we propose SkillCraft, which reformulates post-retrieval selection as a dedicated set-resolution problem. Rather than treating the upstream ranking as an immutable boundary, SkillCraft constructs contextualized list-relative evidence and estimates non-monotonic membership probabilities, dynamically skipping high-ranked distractors to recover non-contiguous skill combinations. Candidate-level membership learning is further coupled with set-level calibration to govern variable-cardinality admission. On the 6600-skill SkillRet benchmark, SkillCraft achieves the strongest skill set-recovery performance and set F1, outperforming recent skill retrieval systems (e.g., SkillRouter, SkillReason, and R3-Skill. Beyond set recovery, its selected skill contexts yield higher downstream task accuracy across three external benchmarks in practice. Moreover, SkillCraft outperforms 17 competitive baselines across sparse, encoder-only, decoder-only, and two-stage retrieve-and-rerank pipelines in candidate retrieval.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.