acceptodds
Under review as a conference paper at ICLR 2027

Defending Against Dormant Poisoning Attacks Across Language and Multimodal Agents

Abstract

Recent work has shown that in dormant poisoning, an attacker implants harmful behaviors that remain hidden at release but emerge later after entirely benign downstream fine-tuning. As foundation models are increasingly used as agents, whether dormant poisoning extends to harmful agent actions and multimodal settings remains unclear. We show two risks of dormant poisoning in agentic settings: (1) activation can extend to tools and actions not optimized during poisoning; and (2) multiple harmful behaviors can coexist within one agent and be jointly activated by fine-tuning. Existing defenses often rely on externally provided safe responses, which are hard to obtain for agent actions, or a trusted clean reference model as a safety anchor. In dormant poisoning, however, the released model appears safe before fine-tuning, so downstream users may be unaware of the poisoning and may not know which model to trust as a clean reference. We address this issue by exploiting a key property of dormant poisoning: the attacker must preserve safe behavior at release to conceal the poisoning, and this preserved behavior can itself serve as a safety reference. Leveraging this property, we automatically construct a Safety Buffer that pairs each harmful input with the released model’s response and a benign version generated by the same model to remove harmful intent. The Safety Buffer can also replace external safe responses required by existing defenses. However, these defenses do not distinguish updates that weaken safety from those that preserve it, often increasing over-refusal or reducing utility. To address this limitation, we lastly introduce LookAhead Defense, which previews each candidate update using the Safety Buffer and penalizes only updates predicted to weaken safe behavior on the harmful input or make the released model’s response more likely on the benign version – allowing us to preserve safety while avoiding unnecessary changes to benign behavior. Across language models, language agents, and multimodal agents, LookAhead Defense reduces harmful behavior in agents from up to 93.9% to at most 14.2% and harmful responses in language models from up to 92.5% to at most 2.0% compared with vanilla fine-tuning, while maintaining competitive utility and low over-refusal. It also remains effective against other poisoning methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.