MalSpot: Detecting and Localizing Malicious Code in LLM Agent Tools
Abstract
Tools extend the capabilities of large language model (LLM) agents, but malicious code in these tools can compromise users' security and privacy. Existing static detectors struggle to distinguish malicious behavior from legitimate functionality because the same operations can serve either purpose. For example, a network request may send credentials to a service for authentication or leak them to an attacker. In this work, we propose MalSpot for statically detecting and localizing malicious behavior in agent tool code. MalSpot first identifies statements that interact with the environment and collects the code that assigns values to the variables they use. With this context, it combines static and LLM-based analysis to select potentially malicious code. Multiple LLMs then independently analyze the selected code together with the tool's complete source code to confirm or revise the initial assessments. MalSpot combines their tool-level decisions by majority vote. For a tool labeled malicious, it reports the location of the statement most frequently identified as malicious. Experiments on three datasets show that MalSpot achieves high true-positive rates with low false-positive rates, outperforming the baselines. With small open-weight models, it achieves performance comparable to direct detection and localization with frontier closed-source LLMs such as Claude-Opus-5 and GPT-5.5.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.