acceptodds
Under review as a conference paper at ICLR 2027

ELLA: Weaponizing Visual Evidence to Hijack Memory Formation in Self-Evolving Agents

Abstract

Persistent memory enables self-evolving agents to summarize interactions and revise state across tasks, but it also allows untrusted evidence to be combined with the agent’s own behavior and persisted as trusted memory. Existing memory-poisoning attacks place a malicious instruction, trigger, or target relation in attacker-controlled content before memory writing. We expose a different attack surface: an attacker can provide only source-side visual evidence and induce the victim’s own memory writer to complete a harmful source–action association during an otherwise ordinary update. We introduce ELLA, which selects benign-looking source evidence that is likely to be retrieved and consumed while leaving the target action and final relation absent from the submitted content. Across multimodal benchmarks, three memory backends, and multiple reader models, ELLA induces persistent targeted behavior after all attacker-owned sources are removed.A no-completion control eliminates the effect despite full source retrieval, and a commitment-flip experiment shows that the activated action follows evidence supplied after the carrier is fixed. These results identify memory formation itself as a security boundary: harmful persistent state can be created by the trusted writer rather than directly injected by the attacker.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.