acceptodds
Under review as a conference paper at ICLR 2027

MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection

Abstract

Agent Skills extend LLM agents with reusable instruction packages that can carry scripts, resources, and service configuration, turning the Skill distribution channel into a direct vector for malicious behavior. Progress on pre-installation detection is limited by its data: existing malicious Skill resources are fragmented across sources, artifact formats, evidence regimes, and benign coverage, and they overlap enough that concatenating published rows both inflates apparent scale and places related content on both sides of any evaluation split. We present MaliciousSkillBench, a consolidated benchmark for malicious Agent Skill detection. From 13 frozen public sources, 11 of which yield eligible malicious Skill artifacts, we recover artifact-level Skills and apply canonicalization, deduplication, structural analysis, and conflict filtering, preserving provenance throughout. The malicious Core reduces from 8,414 raw to 7,539 normalized unique identities organized into 4,588 structural families, yielding 9,740 Skills: 7,505 malicious and 2,235 benign. Using only labels native to each source, we harmonize 11 attack categories for 4,983 malicious identities and find that threat composition differs sharply across sources. On static primary-instruction text, the benchmark reshapes detector conclusions: the strongest sparse lexical baseline falls from 0.932 Random Macro-F1 to 0.665 once whole sources are held out, and a fine-tuned long-context ModernBERT encoder reaches 0.964/0.940/0.734 on Random/structural-disjoint/Source-Disjoint but still carries a 0.230 gap. Class prevalence accounts for little of this gap: evaluation-only standardization removes 5.9% of it, and fully balanced retraining still yields 0.966 vs. 0.662 Macro-F1 and 0.991 vs. 0.718 AUROC. Held-out-source failures are heterogeneous and model-specific: some sources lose ranking quality outright, others rank well but transfer their decision boundary poorly; the encoder over-flags unseen benign sources at 47.3% FPR, whereas a zero-shot LLM judge keeps 0.7% benign FPR but only 41.6% malicious recall. Malicious Skill detection should therefore be reported using three separate quantities: ranking quality, fixed-operating-point performance, and benign false-positive rate, each capturing a distinct failure mode. Code and data: [https://anonymous.4open.science/r/msb-F725/](https://anonymous.4open.science/r/msb-F725/)

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.