Benchmarking and Jailbreaking Phone-use Agents for Social Media Misuse
Abstract
Phone-use agents operate mobile apps through native GUIs, giving them access to functions that social platforms rarely expose through public APIs, such as reporting. Abuse of these functions relies heavily on human labor, and phone-use agents could substantially reduce its cost. Yet existing benchmarks rely on mock apps that offer limited support for abuse relevant functions, and they cover only a narrow range of real-world abuse scenarios. In this work, we take a pioneering step toward systematically evaluating this risk with BadPhoneAgent, a benchmark dedicated to misuse of phone-use agents on social media. It provides a high-fidelity environment spanning 11 social media apps and 18 abuse relevant functions, alongside 150 manually constructed test cases grounded in platform policies and organized into 8 categories and 35 subcategories. Replaying the same tasks on a physical phone reproduces the outcomes observed in our environment (91.5% agreement), indicating that the risks it measures also arise in real apps. To further expose this risk, we propose View Editing for Intent Laundering (VEIL), which exploits the underexplored gap between the model's observation and the phone state. VEIL reaches attack success rates (ASR) of 88% on GPT-6-Astra and 64.7% on Claude Fable 5, while these models refuse only 8.0% and 1.3% of requests, respectively. More concerning, several open-source models surpass these frontier commercial models under VEIL at speed comparable to human baseline. These results reveal the potential for large-scale social media abuse by phone-use agents and provide a foundation for developing safer phone-use agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.