acceptodds
Under review as a conference paper at ICLR 2027

PonderAgent: A Risk-Adaptive Dual-System Framework for Safe and Selective Mobile GUI Agents

Abstract

Mobile GUI agents powered by vision-language models operate in untrusted interfaces, where indirect prompt injection and context-dependent risks can induce unsafe actions. Existing defenses against unsafe or policy-violating agent behavior often rely on model-specific alignment or apply expensive external review uniformly, leading to poor cross-model transferability and high deployment costs. Model-specific alignment is costly to update and may degrade general utility, while uniform review wastes latency, computation, and budget on low-risk inputs. To address these limitations, we present PonderAgent, a risk-adaptive dual-system framework that combines task-level pre-interception, lightweight semantic risk gate, RAG-augmented critic system, and global working memory. Low-risk actions follow a fast path, whereas potentially risky actions are escalated for policy-grounded review, separating task-oriented action generation from safety evaluation. On 230 executable MobileSafetyBench tasks, PonderAgent increases harm prevention for open-weight T-Actors from 13% to 6566%, while goal achievement remains sensitive to T-Actor capability. With GPT-4o as the T-Actor, PonderAgent reaches 70% harm prevention and with Qwen3-VL-32B-Instruct, it prevents 28 of 30 executable indirect prompt-injection attacks (93.3%). In addition, the R-Gate adds 228 ms per step and routes 49.8% of actions to the C-System, reducing the estimated average per-step latency to roughly half the cost of uniform C-System review.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.