CoSA: Characterizing the Conflict of Skills Attack in Real-World LLM Agent Skill Ecosystems
Abstract
An agent skill is a modular, reusable package of instructions and supporting resources that extends large language model agents with composable procedures for complex tasks. Prior work on Skill security has primarily exposed malicious instructions, bundled payloads, and unsafe packages, motivating single Skill package inspection and runtime defenses. However, individually benign Skills may violate system-level or user-defined policies when composed, as their intermediate artifacts and decision signals may be consumed downstream beyond the constraints governing their original use. We formalize such composition induced violations as Conflict of Skills (CoS). We first characterize recurring CoS mechanisms in real-world Skill ecosystems and then introduce the Conflict of Skills Attack (CoSA), a novel mechanism-guided framework for discovering composition opportunities and constructing locally admitted Skill additions. CoSA targets two augmentation strategies arising from Skill composition: (1) capability supplementation, which adds missing operations, and (2) relationship reshaping, which changes how existing capabilities exchange information, propagate decisions, or trigger side effects. We evaluate CoSA on four domains and 288 distinct tasks. Across multiple models and agent harnesses, CoSA raises attack success rates for data leakage, unauthorized state modification, and unsafe execution by about 20 percentage points over the baseline. Trajectory analysis shows that risk arises from cross-Skill information flow, decision reuse, and target/action binding. Based on these findings, we design a lightweight harness-level guard that checks information flow and authorization before consequential actions. In a five-model recruiting pilot, the guard reduces attack success rate by 58.61 percent points while largely preserving benign task completion (66.67 percent points to 65.00 percent points).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.