IntentGuard: Securing GUI Agents without Sacrificing Benign User Intent
Abstract
GUI agents can autonomously interact with webpages, documents, and applications to perform complex user tasks, making them increasingly exposed to untrusted environmental content. Existing work on agent security has begun to defend against untrusted environmental content, but often relies on conservative detection-and-blocking strategies that may cause benign users’ legitimate tasks to be abandoned. When the user request is benign, an effective defense should not only block malicious environmental influence, but also preserve the user’s original intent to complete the legitimate task. To achieve these goals, we propose IntentGuard, a dual-module Blocking-Recovery framework. The Blocking module performs provenance-aware action screening to identify environment-induced unsafe behavior and prevent contaminated context from propagating into harmful actions. The Recovery module generates structured recovery guidance that helps the underlying agent bypass malicious environmental content and resume the user's intended task without retraining. To support systematic evaluation, we introduce EnvGuard-Bench, a GUI-agent safety benchmark organized along two axes: user intent and environmental content, each categorized as benign or malicious, resulting in four interaction settings. We evaluate IntentGuard under these settings, particularly the challenging benign-user/malicious-environment scenario. IntentGuard reduces attack success rate by 63.0 percentage points while achieving 67.0% safe task completion, and further experiments on existing safety benchmarks demonstrate its effectiveness and generalization. Code will be released at: https://anonymous.4open.science/r/IntentGuard-1560
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.