Tracing and Refining Injection Attacks in Agent Memory via Lifecycle Red Teaming
Abstract
Large language model agents increasingly suffer from untrusted data that can contaminate memory and pose major security risks. While recent red teaming research explores various memory injection attacks, current evaluations primarily measure final success rates. They overlook the lifecycle of injections and fail to capture the causal chain of attacks, making it impossible to identify the precise stage where an injection fails. Furthermore, current studies often overlook that injected memories might require interactions with each other to take effect. Tracing these hidden dynamics is difficult in standard systems but becomes exceptionally severe in emerging graph memory environments, where injected inputs undergo complex merging and structural evolution. To address these challenges across diverse platforms, we propose a Lifecycle Red Teaming framework to reconstruct the lifecycle path and refine injected memory attacks. Specifically, we first propose LifecycleTracer to reconstruct fragmented lifecycle paths and expose injection interactions. As the foundation, we establish a formal memory lifecycle chain and a validity gating mechanism to standardize evaluations across diverse memory platforms. To pinpoint precise failure stages for individual injections, we incorporate Intra-Lifecycle Replay Attribution to bridge causal gaps within native logs. To uncover complex memory interactions, we also design Cross-Lifecycle Dynamic Attribution to capture interdependencies among multiple injected inputs. Finally, we leverage this feedback to refine injection attacks. Extensive experiments confirm that our framework effectively traces the lifecycle chain and enhances attack performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.