acceptodds
Under review as a conference paper at ICLR 2027

The Same Pixels, Opposite Authority: Visual Data Injection in Computer-Use Agents

Abstract

Multimodal LLM agents use screenshots to populate tool arguments, but rasterization can preserve a displayed value while omitting the source information needed to authorize its use. We study Visual Data Injection (VDI), in which attacker-controlled visual data supplies a sensitive argument without an injected instruction or a change to the user-authorized operation. We introduce matched-raster, value-counterbalanced pairs that hold the screenshot, user request, and literal value fixed while varying hidden source permissions. This protocol isolates an information boundary: a screenshot-only policy cannot distinguish the two permission states. Under this protocol, Qwen2.5-VL and Qwen3-VL adopt unauthorized candidates in 98/100 observations each, and GPT-4o in 99/100. We propose a defense framework combining independent provenance acquisition, persistent lineage, and role-specific authority enforcement. Reference-origin experiments isolate contract checks using supplied provenance, while browser-origin experiments acquire evidence from live frame metadata. In the latter, enforcement preserves all same-origin candidate executions achieved by the raw models (33/50 for Qwen2.5-VL and 34/50 for GPT-4o) while reducing cross-origin candidate execution to 0/50 for both. Memory experiments show that retaining original provenance prevents unauthorized reuse after storage; error interventions show that unknown provenance triggers confirmation, whereas false trust can permit attacks. These controlled results support separating operation authorization from argument-source authorization, with effectiveness conditional on reliable provenance rather than visual appearance alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.