ScreenLocate: Defending Mobile Agents Against Visual Prompt Injection
Abstract
Mobile agents increasingly rely on screenshots as their only view of an application. This creates a basic security problem because genuine controls and attacker written instructions occupy the same visual space. An instruction placed in a message, review, or pop up can therefore redirect the next action without changing the agent or the application. Existing defenses usually either convert the screen into text, which discards layout, or alter the victim model, which limits their use with closed systems. We introduce \ScreenLocate, a defense that uses the trusted task to find visual regions likely to redirect the agent and removes only those regions before inference. \Lite provides a lexical version of this idea, while \Ours also uses screen geometry, repairs directives split across adjacent text regions, and fills selected areas from their local boundaries. On the scenario disjoint \AgentHazard test, \Ours lowers attack success from – to – across three multimodal agents with a change of at most points in benign task performance. Under nine semantic presentation mutations, the fixed lexical variant lowers attack success by – points across two primary victims and an additional Qwen3-VL-8B victim. Our experiments distinguish blocking the attacker from recovering the intended task, which reveals when a seemingly successful defense merely causes an invalid action. The results support a simple principle. Visual prompt injection should be handled as a localized intervention whose success is evaluated through the action taken by the agent.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.