BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning
Abstract
LLM agents increasingly rely on installable skills that package task logic, schemas, runtime code, and sometimes learned components. This creates a supply-chain blind spot: source review can reveal conditional routing, while bundled weights still determine which structured invocations activate it. We present BadSkill, a model-in-skill backdoor whose target condition is an attacker-chosen conjunction of schema-valid parameter assignments. Each assignment remains interface-valid, while the target label depends on their joint structured context rather than an isolated anomalous token. During training, full-trigger positives and independently generated near-trigger hard negatives teach a bundled classifier an approximate activation boundary. At runtime, a source-visible router conditionally executes a harmless canary side effect while returning the advertised task result in either branch. We evaluate eight target skills, five non-trigger controls, eight models, and four integration settings. Across 32 model-integration cells, BadSkill attains 97.2% mean canary-delivery ASR, while 97.5% of non-triggered runs return a visible result, at most 4.5 points below clean artifacts. The two rates separately characterize trigger-side delivery and non-triggered result availability under the reported protocol. Activation remains high across model scales and clients but is more sensitive than visible-result return to character-level corruption. The evaluation motivates treating bundled models and their execution behavior, not only source and documentation, as part of the installation trust boundary.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.