From Declared Capabilities to Behavior: Detecting Malicious Agent Skills Through Consistency Analysis
Abstract
A skill is a natural language document, often bundled with scripts, that an LLM agent loads on demand. Malicious skills can abuse an agent’s privileges to exfiltrate sensitive information, compromise the host environment, and influence subsequent agent actions. Detecting such skills before deployment is therefore critical. To enable scalable security vetting, researchers have explored both dynamic and static analysis. Dynamic analysis requires executing potentially relevant behaviors, resulting in substantial time and token costs that make it difficult to scale to large skill collections. Static analysis is more scalable because it can identify suspicious operations without execution, but it often produces false positives because security sensitive operations may also be necessary for benign skills. To improve detection, we compare a skill’s declared capabilities with its actual behavior. SKILL.md specifies the declared capabilities, while our analysis identifies security-sensitive behaviors across code, files, and document instructions. Building on this comparison, we introduce a declaration-behavior consistency analysis that detects mismatches between declared and actual capabilities and reports the associated source, sink, and declaration. We evaluate our method on 5,520 skills and show that it outperforms five static-analysis baselines, achieving an F1 score of 0.946 and an average precision of 0.984. It also maintains 0.86 recall under the strongest evasion family and 0.90 macro-averaged recall on previously unseen attack types.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.