acceptodds
Under review as a conference paper at ICLR 2027

Learning Operational Boundaries for GUI Agents via Constrained Optimization

Abstract

While graphical user interface (GUI) agents have made significant strides in executing complex tasks across diverse environments, their interaction with open-ended web interfaces raises substantial safety risks. Adversarial or compromised websites deliberately designed to mislead users and agents into unsafe operations pose a common and severe threat. In this work, we introduce GUI-WebMinefield, a benchmark for evaluating GUI agents in harmful web environments. Our evaluation shows that even state-of-the-art agents often fail to recognize potential risks and still attempt to execute instructions in malicious settings. To address this gap, we propose Constraint-Conditioned Optimization with Operation Advantage Calibration, which improves risk-aware action selection by incorporating explicit safety constraints. We construct safety constraints to define permissible operational boundaries in GUI environments and perform token-level calibration of operation-specific advantages during reinforcement learning. Experiments demonstrate that our method preserves strong task performance while substantially improving safety, without requiring exposure to harmful environments during training. The benchmark and code will be made publicly available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.