SafeDroid: The Cost of Obedience for VLM-Based Mobile Agents
Abstract
Mobile agents built on vision-language models (VLMs) operate a phone on the user's behalf, changing messages, accounts, and settings. Based on a questionnaire of 2,240 smartphone users, we found that 11% of requested phone changes were judged harmful in context by independent participants. However, agent-safety research has concentrated its threat model on externally injected goals—prompt injection, manipulated interfaces, untrusted tool output—so on mobile devices the user's own request is rarely the object of evaluation. We distill the 11% into twelve harm domains and introduce SafeDroid, a survey-grounded stress test of mobile agents' responses to risk-sensitive user requests, with device-state verification. Evaluating eleven general-purpose VLMs and five specialized GUI models reveals severe vulnerabilities: as single agents, the general-purpose VLMs reach the requested device states (target states) in 58.3% of runs without any external manipulation, and 93.6% of those trajectories never identify the specific consequence. Although Mobile-Agent-v3 (MA-v3) lowers the overall target-state reach rate, this reduction stems primarily from capability failure rather than cautious refusal. We call this the cost of obedience: overlooking consequences while faithfully carrying out a user request. Therefore, mobile-agent safety must assess the consequences of user requests even when they arrive through a trusted input channel.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.