Parallax: A Two-Phase Defense for Memory-Augmented LLM Agents from a Causal Perspective
Abstract
Memory poisoning allows adversaries to implant persistent records that steer LLM agents toward unauthorized actions across future queries. Content-level filters struggle with benign-looking, cooperative records, while action-level checks can block execution without removing the memory source of the threat. We propose PARALLAX, a defense that connects online interception with offline attribution through a causal perspective. Online, PARALLAX constructs counterfactual memories along the epistemic, procedural, and episodic pathways, aiming to suppress malicious influence while preserving beneficial task-relevant content. It compares actions under original and counterfactual memories in synchronized execution contexts, blocking actions that differ substantively and lack support from the user query, then prompting replanning. Offline, we define minimal trigger sets (MTSs) based on causal sufficiency and necessity for triggering an intercepted action. To accelerate the search for MTSs, PARALLAX generates candidate sets using an LLM guided by replay feedback and prior attribution experience, verifies them through repeated replay, and removes verified sets from memory. Across three agents and three adaptive memory-poisoning attacks, PARALLAX achieves the lowest attack success rate (0% on EHRAgent and StateAct) and 83.3–100.0% attribution recall at 100.0% precision, while preserving benign utility. These results demonstrate the advantages of a causal perspective for memory-poisoning defense.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.