ELLA: Weaponizing Visual Evidence to Hijack Memory Formation in Self-Evolving Agents
Abstract
Persistent memory enables self-evolving agents to summarize interactions and revise state across tasks, but it also allows untrusted evidence to be combined with the agent’s own behavior and persisted as trusted memory. Existing memory-poisoning attacks place a malicious instruction, trigger, or target relation in attacker-controlled content before memory writing. We expose a different attack surface: an attacker can provide only source-side visual evidence and induce the victim’s own memory writer to complete a harmful source–action association during an otherwise ordinary update. We introduce ELLA, which selects benign-looking source evidence that is likely to be retrieved and consumed while leaving the target action and final relation absent from the submitted content. Across multimodal benchmarks, three memory backends, and multiple reader models, ELLA induces persistent targeted behavior after all attacker-owned sources are removed.A no-completion control eliminates the effect despite full source retrieval, and a commitment-flip experiment shows that the activated action follows evidence supplied after the carrier is fixed. These results identify memory formation itself as a security boundary: harmful persistent state can be created by the trusted writer rather than directly injected by the attacker.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.