acceptodds
Under review as a conference paper at ICLR 2027

PRISM: Probe-Guided Localization and Feature Repair for Visual Prompt Injection

Abstract

A visual prompt injection defense can suppress the attacker's target answer yet still fail the user's task. It must reject image-borne instructions while preserving the visual evidence needed to answer correctly. Separate-caption evaluations do not test this distinction within application content. We introduce AppInjectBench, a benchmark spanning 11 applications that embeds injected instructions in interface content such as reviews, comments, and messages. Each example pairs a clean screenshot with a benign-text variant and an injected variant, enabling joint evaluation of task recovery and benign task performance. We further propose PRISM, which observes attention to visual tokens under a fixed query. A learned spatial locator uses this observation to identify injection regions. PRISM replaces the selected visual features with a shared mean of their unselected boundary features. The frozen model then answers the original task from a fresh prefill using the modified representation. On model-specific AppInjectBench test sets of 1,000 sources screened for correct benign answers and successful undefended injection, PRISM recovers 96.7% of cases with Qwen2-VL-7B and 94.0% with Qwen2.5-VL-7B, while causing new errors on at most 0.3% of benign inputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.