LatentTriad: Pre-Action Representation Probing for Malicious Agent Skill Detection
Abstract
Large language model agents are increasingly extended with reusable skills. Agent skills are becoming a key interface for extending agent capabilities, but also introduce new attack surfaces that enable security violations. Detecting malicious skills at ecosystem scale remains challenging. Existing approaches broadly fall into static and dynamic detection. Keyword-based static methods can incur high false-positive rates. Dynamic methods require costly execution, limiting scalability to large skill repositories (e.g., SkillsMP indexes over 3.18 million public SKILL.md files). This creates a key challenge: capturing the true malicious behavior while remaining lightweight and scalable. We introduce LatentTriad. Rather than relying on string-pattern matching, our key insight is to probe the model's hidden representations before the output tokens are executed. Specifically, we find that whether an operation is consistent with the trusted task is encoded in the model's hidden states, and this signal generalizes across held-out expression forms in controlled evaluations. Importantly, LatentTriad operates only during the model's prefill stage and does not generate output tokens, avoiding the skill execution required by execution-based analysis. We evaluate LatentTriad on 5,830 skill packages from five source datasets. Across seven backbones, LatentTriad achieves 99.0–99.2% pooled F1. Our compact implementation completes prefill and detection in a median of 259–284 ms, whereas the sandbox-based SkillDetonate reports approximately 153 s per skill, a roughly 540–590-fold difference in reported latency. Controlled interventions further show that replacing operation states with those from matched examples of opposite risk shifts fixed risk scores toward the replacement examples' labels. These results demonstrate that LatentTriad is practical and scalable for detecting malicious skills.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.