Spoofing the Toolbox: Optimizing Tool Descriptions to Deceive LLM Agents
Abstract
Large language model (LLM) agents are widely used in many applications and often rely on external tools. Tool information, such as function descriptions, helps agents determine which tools best suit a given context. However, attackers can manipulate this information to mislead LLM agents into using malicious tools. Prior attacks mainly rely on keyword-based techniques, which are easily detected, or rule-based methods, which have limited effectiveness. In this work, we introduce TSA, a stealthy tool-spoofing attack that crafts descriptions for malicious tools that are difficult to distinguish from those of existing benign tools. To achieve this, we propose an LLM-powered optimization method that iteratively refines malicious tool descriptions. Specifically, TSA begins with a variant of a benign tool description and updates it based on feedback from a surrogate LLM agent. The optimization method incorporates a feedback function and contrastive in-context learning to guide the search for descriptions that are both stealthy and effective. Experiments on both single- and multi-tool use benchmarks demonstrate that TSA achieves an attack success rate of up to 92.13%, outperforming four state-of-the-art approaches while remaining highly stealthy. Our user study further shows that even human participants cannot distinguish descriptions generated by TSA from benign ones. Moreover, TSA remains effective against five tool-malware detectors.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.