acceptodds
Under review as a conference paper at ICLR 2027

BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning

Abstract

LLM agents increasingly rely on installable skills that package task logic, schemas, runtime code, and sometimes learned components. This creates a supply-chain blind spot: source review can reveal conditional routing, while bundled weights still determine which structured invocations activate it. We present BadSkill, a model-in-skill backdoor whose target condition is an attacker-chosen conjunction of schema-valid parameter assignments. Each assignment remains interface-valid, while the target label depends on their joint structured context rather than an isolated anomalous token. During training, full-trigger positives and independently generated near-trigger hard negatives teach a bundled classifier an approximate activation boundary. At runtime, a source-visible router conditionally executes a harmless canary side effect while returning the advertised task result in either branch. We evaluate eight target skills, five non-trigger controls, eight models, and four integration settings. Across 32 model-integration cells, BadSkill attains 97.2% mean canary-delivery ASR, while 97.5% of non-triggered runs return a visible result, at most 4.5 points below clean artifacts. The two rates separately characterize trigger-side delivery and non-triggered result availability under the reported protocol. Activation remains high across model scales and clients but is more sensitive than visible-result return to character-level corruption. The evaluation motivates treating bundled models and their execution behavior, not only source and documentation, as part of the installation trust boundary.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.