acceptodds
Under review as a conference paper at ICLR 2027

Reality Flow Integrity: Defending VLM-Enabled Wearable Assistants against Physical Prompt Injection

Abstract

VLM-enabled wearable assistants continuously observe physical environments, where task-relevant evidence and attacker-controlled text enter through the same camera frame. This shared visual channel provides no explicit trust boundary and enables physical prompt injection, which allows attacker-controlled visual text to redirect the assistant's response. A naive mitigation against physical prompt injection is to mask all visual text before it reaches the VLM. However, this mitigation removes information needed for core tasks, such as reading signs and documents, and therefore reduces utility. To address this problem, this paper presents RFI, which creates a visual trust boundary within each frame and uses a dual-VLM architecture so that the trusted VLM can reason about the scene without reading untrusted visual text. RFI detects and masks regions that contain text, separating them from the rest of the camera frame. It then assigns different inputs and roles to two VLMs, a trusted VLM (T-VLM) and an untrusted VLM (U-VLM). The T-VLM reasons over the masked frame and answers the user's request, while the U-VLM reads text from isolated regions and returns only opaque symbols to the T-VLM. This separation prevents the T-VLM from following malicious instructions while preserving utility for tasks such as reading signs and documents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.