acceptodds
Under review as a conference paper at ICLR 2027

Preserving the Task: What Training Against Poisoned Memory Learns

Abstract

Memory poisoning can redirect LLM agents through malicious instructions in retrieved records. To defend against these attacks, the agent must resist those instructions without discarding useful facts or preferences that are compatible with the user's task. We introduce task-preservation training, a supervised fine-tuning approach that aims to resist memory poisoning while preserving the use of benign content. We use the model's own responses to clean inputs, filtered for correct final answers, as targets under conflicting records and include examples that follow a benign formatting instruction. Training lowers attack success on memory-poisoning benchmarks and transfers to tool-calling tasks in AgentDojo. On Qwen2.5-32B, attack success falls from 16.4% to 2.1% and tasks completed without attack success rise from 52.8% to 64.3%, at some cost in clean utility. We then examine which uses of memory the trained model retains. On a smaller Qwen2.5-14B, it still answers from retrieved facts when we edit them to change the correct answer, and it still follows formatting instructions from memory. We find that it resists instructions that would replace its answer, including unseen replacement attacks, but only partially resists attacker content appended to a correct answer. With activation patching, we locate this resistance in how the trained model represents the attacker's demand within the record at intermediate layers. Overall, our study shows transferable protection from narrow supervision while distinguishing attack resistance from the ability to use memory reliably as the task requires.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.