DeepSift: Attributing Malicious Tool Invocations in LLM Agents
Abstract
An LLM agent augments a backbone LLM with tools and memory, and its context accumulates the user query, memory records, demonstrations, and tool observations while it acts. By injecting malicious texts into one of these inputs, for instance a poisoned memory record or an injected prompt, an attacker can drive the agent to invoke a tool of the attacker's choice, and existing defenses cannot prevent all such attacks. Once such an attack succeeds, identifying the texts in the context that are responsible for the malicious tool invocation is crucial for post-attack analysis and mitigation, a task we call attribution for LLM agents. Existing attribution methods score the entire context at once or target one type of input, so they are inaccurate for LLM agents: unrelated context dilutes the contribution of the malicious texts, and content generated after the attack masks it. To this end, we propose DeepSift, which organizes the context as a segmentation tree, recursively prunes it by Shapley scores to counter dilution, and excludes later-generated content to avoid masking, yielding a compact set of candidate texts. DeepSift then performs interaction-aware Shapley attribution over these candidates, so that each fragment of a split malicious text is still identified even if it is ineffective alone, and adaptively selects the malicious texts at the largest score gap. Theoretically, we bound how much the malicious texts distort the scores of benign texts and give conditions under which DeepSift returns exactly the malicious texts. Extensive experiments across diverse agent scenarios, attack families, and backbone LLMs show that DeepSift accurately attributes the malicious texts and outperforms existing methods, even under strong adaptive attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.