acceptodds
Under review as a conference paper at ICLR 2027

AutoSkillDrain: Resource Amplification Attacks on LLM Agents via Malicious Skill Optimization

Abstract

LLM agents increasingly rely on reusable skills to acquire domain-specific capabilities, but this extensibility also creates a new supply-chain surface for resource-amplification attacks. A malicious skill can induce prolonged generation, thereby increasing cost, latency, and service pressure. To systematically study this threat, we propose AutoSkillDrain, the first black-box resource-amplification attack framework that automatically transforms benign agent skills into resource-draining variants without access to the victim agent's system prompts, or internal execution policies. We further introduce SkillDrainBench, a benchmark consisting of public source skills, optimized malicious counterparts, and task instances for evaluating amplification, utility, transferability, and detectability. Across experiments on three agent backends and multiple models, AutoSkillDrain achieves the strongest resource amplification among compared methods, with 2.57, 2.15, and 2.28 average output amplification on PydanticAI, OpenClaw, and Claude Code, respectively, while maintaining over 91% task completion. These results expose an underexplored resource-security risk in open skill ecosystems and highlight the need for validation before deployment.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.