acceptodds
Under review as a conference paper at ICLR 2027

What Did the Agent See? Implanting False Memories in Multimodal AI Agents through Black-Box Visual Attacks

Abstract

Multimodal AI agents increasingly rely on persistent long-term memory to ground generation in past visual episodes, and in embodied deployments that memory becomes the agent's record of an environment it can no longer observe. We show that unconditional trust in stored visual content creates a critical vulnerability: an attacker who controls nothing but an image can make an agent remember something that never happened. We study a strictly image-bounded threat model: a content publisher perturbs an image a user later shares with their agent, with no access to the target model, its memory system, the text channel, or future queries. The poison must prevail over the user's truthful text and remain dormant until a dependent query retrieves it. We propose Vivid, a blind, black-box framework on two axes: scene-dominant poison design, which constructs from the image alone the content that survives transfer through pooled embeddings, frozen before any query exists; and delivery-aware optimization, which aligns the image with that target across encoder families and in expectation over the re-encoding it undergoes before ingestion. Across four benchmarks, five memory backends and four black-box victims, Vivid succeeds against every backend and outperforms the strongest prior transfer attack, decisively so under realistic delivery channels. We further characterize purification and memory-level detection and isolation as defenses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.