Grounding Is Not Authorization: Teaching GUI Agents When Not to Act
Abstract
GUI agents increasingly act on behalf of users, yet a screenshot shows which action is available, not whether the user asked for it. We find that agents fail to withhold in two ways: a task-relevant screen suppresses clarification of under-specified requests, and neutral wording elicits high-impact actions that harm-framed wording does not, although neither request authorizes the operation. Agents thus act unless a cue stops them, rather than by execution justification: whether the request itself warrants the action. We introduce AXIS, a protocol that crosses request phrasings with screen conditions while fixing the operation and never supplying a warrant. Across five models, a matching screen cuts clarification from 60–73% to 3–21%, and neutral wording raises execution by up to 72 points; safety prompts and inference-time checks do not close the gap. We propose AXIS-Core, which trains a GUI agent to assess operation risk and request authorization before acting. On UI-Venus-1.5-8B, it makes no unsafe action on 134 danger records while completing all 18 required benign actions, and separates 49 of 50 authorization-contrast pairs where untrained methods separate none. It generalizes to unseen phrasings and screenshot sources with 92–97% accuracy, and keeps native grounding and navigation within 1.7 points of the base model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.