EvoAttack: Adaptive Strategy Learning for Agent Attacks with Evolving Memory
Abstract
Memory allows LLM agents to carry information across interactions, making their behavior increasingly shaped by an evolving internal state. This same capability makes attacking memory based agents fundamentally nonstationary, since the effectiveness of an attack strategy can change with both the target memory state and subsequent interactions. Existing memory attacks primarily focus on injecting or sustaining malicious information, while adaptive attack methods refine attacks from local feedback, but neither explicitly learns how strategy utility evolves with the target. We propose , which formulates attacks on memory based agents as adaptive strategy learning against an evolving target under black box access. EvoAttack infers the target state from interaction feedback and assigns temporal credit using attack outcomes observed over multiple horizons. It then maintains an attacker memory that links target states, strategies, and realized rewards, turning previously isolated attack outcomes into reusable utility knowledge. Retrieved memories serve as a nonparametric critic to estimate state conditioned strategy advantages and directly recalibrate the LLM attack policy without parameter training. Extensive experiments across diverse agent domains, model backbones, initial memories, and defensive settings show that EvoAttack consistently improves attack effectiveness across both target states and evaluation horizons. These results suggest that effective attacks on memory-based agents require learning how strategy utility changes with the evolving target, rather than repeatedly optimizing against a fixed target condition.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.