acceptodds
Under review as a conference paper at ICLR 2027

MELT:A Dependency-Aware Defense against Memory Poisoning Attacks Based on Taint Analysis

Abstract

Self-evolving LLM agents update their internal state across sessions, often by storing and reusing long-term memory. This improves performance on long-horizon tasks but also introduces a security risk: attackers can poison long-term memory and allow malicious influence to persist during later retrieval and reasoning. More challenging, poisoned memories may not contain explicit malicious instructions. Malicious intent can be hidden in seemingly plausible facts, relations, or reasoning fragments, or distributed across multiple memories that appear harmless when examined individually. As a result, detection methods that analyze each complete memory as a single unit may fail to identify the resulting security risks. To address this problem, we propose MELT (Memory Exposure and Latent Taint), a fine-grained security analysis framework for memory-augmented LLM agents. MELT decomposes natural-language memories into dependency relations and organizes them in a Memory Dependency Graph (MDG). The MDG captures dependencies both within individual memories and across multiple memories. MELT then performs relation-level taint propagation to track how untrusted information spreads along these dependencies. Finally, it validates tainted relations against independent evidence to identify dependencies that may affect security-sensitive reasoning or behavior. Across 298 evaluation units on the native AgentDojo architecture, MELT detects 81.3% of unauthorized-access attacks and 83.9% of log- or fact-tampering attacks. It also intercepts 66.7% of compositional attacks at retrieval time, with a 0.0% false-positive rate on benign cases. These results show that MELT can identify harmful dependencies hidden in seemingly plausible memories, including attack paths that emerge only when multiple memories are combined.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.