Preserving the Task: What Training Against Poisoned Memory Learns
Abstract
Memory poisoning can redirect LLM agents through malicious instructions in retrieved records. To defend against these attacks, the agent must resist those instructions without discarding useful facts or preferences that are compatible with the user's task. We introduce task-preservation training, a supervised fine-tuning approach that aims to resist memory poisoning while preserving the use of benign content. We use the model's own responses to clean inputs, filtered for correct final answers, as targets under conflicting records and include examples that follow a benign formatting instruction. Training lowers attack success on memory-poisoning benchmarks and transfers to tool-calling tasks in AgentDojo. On Qwen2.5-32B, attack success falls from 16.4% to 2.1% and tasks completed without attack success rise from 52.8% to 64.3%, at some cost in clean utility. We then examine which uses of memory the trained model retains. On a smaller Qwen2.5-14B, it still answers from retrieved facts when we edit them to change the correct answer, and it still follows formatting instructions from memory. We find that it resists instructions that would replace its answer, including unseen replacement attacks, but only partially resists attacker content appended to a correct answer. With activation patching, we locate this resistance in how the trained model represents the attacker's demand within the record at intermediate layers. Overall, our study shows transferable protection from narrow supervision while distinguishing attack resistance from the ability to use memory reliably as the task requires.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.