CARVE: Coverage-Aware Reconciliation with Verified Evidence for Agent Skill Detection
Abstract
Agent skills combine natural-language instructions, executable code, and declared capabilities. Their attacks may target the instruction layer, the code layer, or both, creating a heterogeneous detection problem. Static scanners are safer to deploy but provide incomplete coverage when attacks fall outside their rules or observed signals. Dynamic scanners broaden behavioral coverage by executing skills, but require sandboxing, incur substantial operational cost, and remain exposed to sandbox-escape risk. Real deployments also impose confidentiality and intellectual-property constraints: sending proprietary skills to proprietary LLM may disclose sensitive content, motivating locally deployable open-weight analysis. We present CARVE (Coverage-Aware Reconciliation with Verified Evidence), a non-executing detector that separates evidence localization from judgment: code screening and evidence-guided investigation analyze implementation behavior, while independent full-text checks assess instruction injection and potential consequences. Typed signals are fused conservatively, source-linked evidence and channel provenance are retained for audit, and governance findings are reported separately. We evaluate CARVE with open-weight models spanning 7B, 72B, and approximately 744B total parameters, as well as two closed-source models. Across MalSkillBench and MaliciousSkillBench, CARVE maintains strong detection performance across different model scales. Ablations show that each of the three analysis modules contributes to detection performance, with semantic checks and evidence-guided investigation providing the largest coverage gains. A separate cost evaluation on MaliciousSkillBench shows that multi-view analysis reaches a higher-recall operating region. The results support coverage-oriented, evidence-grounded assessment for deployable open-weight models, while highlighting the remaining limits of static analysis and model-dependent judgment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.