Detecting Malicious Agent Skills with Learned Behavioral Rules
Abstract
While agent skills extend the capabilities of large language model (LLM) agents, they may carry malicious instructions that cause agents to perform harmful actions. Existing detectors do not learn interpretable detection criteria from data, which limits their coverage and interpretability. We propose Skill Behavioral Rules (SBRs), a class of interpretable behavioral rules learned from skill graphs for malicious skill detection. We construct a behavior graph from each skill document, augment the graph with machine learning (ML) predicates that score the maliciousness of textual descriptions, and mine rules in Horn form that combine structural patterns with these predicates. For detection, we learn rule weights with a sparse logistic regression model and make each prediction traceable to the matched rules and their supporting evidence. On 1329 MalSkillBench test skills, SBRs outperform nine baseline configurations with an F1 score of 87.17% and a 5.88% false positive rate. The learned rules also generalize to ASM, a dataset from a different source, which indicates that the mined patterns capture transferable behavioral characteristics. Our code and data are available at https://anonymous.4open.science/r/SBR-52F1/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.