Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
Abstract
Persistent memory makes agent safety a longitudinal property: an ordinary request can become unsafe because of state accumulated much earlier. Measuring this effect is difficult because cumulative failure counts also change when the interaction stream itself changes. We introduce a trigger–probe protocol that evaluates fixed benign requests against persistent states reconstructed at increasing exposure levels, with matched memory-free counterfactuals for operational attribution. Across configurable memory architectures, three model backbones, and synthetic and real correspondence streams, memory-induced violation rates generally increase with accumulated experience. In persistent OpenClaw workspaces, we additionally distinguish retained state from later outcomes and find that sensitive values can continue spreading across workspace artifacts even after the set of retained values has stabilized. This decomposition exposes an observable intervention point for retrieval-mediated agents: the context selected from memory is available before generation. We package the longitudinal evaluation workflow as MemVax.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.