acceptodds
Under review as a conference paper at ICLR 2027

Detecting Malicious Agent Skills with Learned Behavioral Rules

Abstract

While agent skills extend the capabilities of large language model (LLM) agents, they may carry malicious instructions that cause agents to perform harmful actions. Existing detectors do not learn interpretable detection criteria from data, which limits their coverage and interpretability. We propose Skill Behavioral Rules (SBRs), a class of interpretable behavioral rules learned from skill graphs for malicious skill detection. We construct a behavior graph from each skill document, augment the graph with machine learning (ML) predicates that score the maliciousness of textual descriptions, and mine rules in Horn form that combine structural patterns with these predicates. For detection, we learn rule weights with a sparse logistic regression model and make each prediction traceable to the matched rules and their supporting evidence. On 1329 MalSkillBench test skills, SBRs outperform nine baseline configurations with an F1 score of 87.17% and a 5.88% false positive rate. The learned rules also generalize to ASM, a dataset from a different source, which indicates that the mined patterns capture transferable behavioral characteristics. Our code and data are available at https://anonymous.4open.science/r/SBR-52F1/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.