VPIXDistill: Securing Visual Inputs against Prompt Injection via Explainability-Guided Distillation
Abstract
Large vision-language models (LVLMs) increasingly use visual text as an important source of task evidence across document, web, and physical-world applications. However, this strength also exposes LVLMs to text-form visual prompt injection (VPI), where malicious text in external visual content can bias the model toward an attacker-chosen answer or hijack it into following an attacker-specified instruction. A practical defense requires more than mitigating attacks. It must suppress malicious visual text while preserving the model's visual-text understanding and reasoning abilities. To make this requirement measurable, we introduce SafeRegionVPI, a bounding-box (bbox) level benchmark with explicit annotations for both injected text and benign text. SafeRegionVPI tests whether a defense distinguishes injected text from benign visual evidence while preserving visual-text understanding beyond a single task-level answer. We further propose VPIXDistill, an explainability-guided defense for visual prompt injection in LVLMs. VPIXDistill uses the LVLM's internal signals to distinguish injected visual text from benign visual evidence and distills them into a lightweight defense model for efficient inference time protection. We evaluate VPIXDistill on SafeRegionVPI, representative VPI benchmarks, and case studies on web agent hijacking and drone task manipulation. VPIXDistill reduces attack impact while preserving model utility with low inference overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.