MobileGrail: Learning Compact Guardrails for On-Device deployable Mobile GUI Agent Safety
Abstract
Runtime guardrails protect mobile GUI agents by assessing each proposed action before execution, but existing guardrails rely on cloud hosted large multimodal models, exposing every screen to a remote server, including private information such as payments, messages, and medical records. We therefore study how to train a compact guardrail that runs on device, which raises two challenges. On-device budgets limit the model to a short context, yet safety decisions hinge on evidence observed many steps earlier, which existing guardrails lose in truncating or summarizing the history. A compact model also fails to judge safety from a single prompt, so its safety boundary must be learned from labels, yet existing labels mark anticipated risk and committed violations alike as unsafe, which blurs the boundary and rejects benign tasks that share the same risk indicators. We propose MobileGrail, an on-device deployable model that addresses both. It maintains a recurrent task state that retains only the evidence subsequent safety decisions depend on, so that authorization and execution progress persist beyond the short context. To intervene early while keeping the safety boundary precise, it learns two labels, unsafe and block: unsafe marks the action that commits a violation, and block is admissible over an intervention window that opens once the evidence justifies intervention. To supply the learning signal, we build a training environment of 32 applications whose backends verify the effect of every action, grounding both labels and providing verifiable rewards with which reinforcement learning calibrates intervention timing. We further build MobileVeto, a benchmark of held-out applications annotated with intervention windows and paired with benign counterparts. On MobileVeto, MobileGrail built on Qwen3.5-2B raises effective intervention over its prompted backbone from 64.5% to 76.2% and reduces premature intervention from 58.6% to 5.2%, outperforming OS-Sentinel with Kimi-K2.6 by 36.9 and 41.4 points in effective intervention and benign false positives. Quantized to 4 bits, it runs on a Snapdragon 8 Elite at a median latency of 4.7 s, comparable to cloud-hosted Gemini 3.8 Flash (5.1 s).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.