acceptodds
Under review as a conference paper at ICLR 2027

MEMSHIELD: NEUTRALIZING COALITIONAL MEMORY POISONING IN LLM AGENTS VIA CAUSAL REPRESENTATION REPAIR

Abstract

Persistent memory empowers LLM agents but exposes them to joint memory-poisoning attacks that manipulate the retrieval representation space to induce cross-session harmful responses. While existing defenses effectively detect isolated poisoned records, they fall short against multi-memory attacks that distribute adversarial triggers across records. To bridge this research gap, we present , a causal memory-repair framework that retrieves and cleanses combined poisoned memories. We formulate memory repair as a causal representation intervention problem and introduce a multi-agent planner that determines optimal memory subsets to retain, mask, or quarantine. Specifically, the planner isolates the Minimal Triggering Set (MTS) driving the risk and the Minimal Blocking Set (MBS) required for risk suppression. To ensure repaired memory spaces cannot reactivate adversarial circuits, incorporates a closed-loop verification mechanism with pre-retrieval quarantine and post-retrieval re-gating. Across four attack families and six benchmarks, reduces the residual Attack Success Rate (ASR) to 0% with negligible impact on clean performance, while demonstrating strong cross-architecture transferability.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.