Implicit-Tool Attacks in Coding Agents
Abstract
Coding agents are increasingly integrated into software development workflows. Yet existing safeguards for coding agents primarily scrutinize their declared tool interfaces and may miss the consequences of persistent configuration edits consumed by downstream components. We term attacks exploiting this gap **implicit-tool attacks**: an adversary induces an agent to configure a benign, unmodified component to perform a sensitive operation later, outside the agent's direct execution path. To systematically evaluate this threat, we design ImpCraft, an automated framework that discovers implicit capabilities in benign extensions, synthesizes scenario-consistent attack instances, and validates their end-to-end runtime effects. Across 810 trials spanning 21 real-world extensions, three commercial products and three frontier foundation models over three independent rounds, implicit-tool attacks achieve an overall success rate of 80.1%, compared with 0.2% for matched direct requests seeking the same effects. Motivated by the role of downstream consequence recognition in attack failures, we propose a consequence-guided reasoning defense that directs agents to inspect referenced programs, endpoints, and execution paths before modifying configuration. In a focused GPT-5.5 evaluation across the three products, this defense substantially reduces attack success. This work provides insights for coding-agent security governance by uncovering an attack surface across component boundaries.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.