acceptodds
Under review as a conference paper at ICLR 2027

GUI-ACE: ACTION-CONDITIONED EXPLANATION OF GUI INTERACTIONS

Abstract

To serve as practical personal assistants, graphical user interface (GUI) agents need to execute tasks reliably and provide personalized services. Interaction records can support large-scale agent training and user behavior analysis, and their reuse benefits from understanding recorded actions' targets and functions. General vision-language models (VLMs) can describe these actions, but large models are costly at scale. They also struggle to align numerical action coordinates with visual user interface (UI) elements. Even with accurate point localization, they may misidentify the complete functional unit receiving the action. Element boxes or an additional screenshot can aid target identification but may be unavailable or add overhead. We therefore introduce ActionScope, the smallest complete UI region corresponding to the action-receiving functional unit. This region specifies which functional unit the explanation should describe. Building on ActionScope, we propose Action-Conditioned Explanation (ACE), a compact VLM that takes a single screenshot, action parameters, and available previous-step textual context. It generates target and purpose descriptions conditioned on its predicted ActionScope bounding box field. Our ACE-Curator pipeline generates ActionScope and semantic annotations through region guidance and quality control. We use these annotations to train 1.3B and 8B variants. On an evaluation set with human-verified reference annotations, ACE-1.3B and ACE-8B achieve 92.50% and 94.54% joint semantic accuracy, respectively. Single-screenshot ACE-8B outperforms Qwen3.8-2.4T (Max API) and Claude Opus 5 with an additional screenshot in overall joint semantic accuracy. Ablations show that ActionScope supervision improves joint semantic accuracy. Replacing predicted boxes with incorrect ones reduces subsequent explanation accuracy. Together, these findings support ActionScope as an intermediate representation connecting spatial evidence to semantic explanation. For offline GUI navigation and grounding, we train models with instructions reconstructed from ACE descriptions. These models outperform those trained with original instructions on several benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.