acceptodds
Under review as a conference paper at ICLR 2027

Reframing GUI Grounding as Intent-Level Object Detection

Abstract

Currently, Vision-Language Model (VLM)-based GUI agents formulate grounding as coordinate prediction, forcing the model to learn a mapping from spatial locations to discrete coordinate tokens without explicit spatial supervision. To address this limitation, we revisit GUI grounding from a spatial detection perspective and find that existing object detectors propose GUI-element-like regions but rarely rank the intended one first, i.e., they lack high-level intent understanding. We therefore propose IntentLoc, an intent-level object detection model that uses high-level semantic queries to equip a detector with intent understanding ability and introduces Crosshair Attention to progressively locate small GUI targets. Under the same backbone and training data, IntentLoc-GUI outperforms pure coordinate prediction baselines by 10.6 points, averaged with equal weight over four GUI grounding benchmarks and two model sizes. We further introduce Atypical-GUI-Bench, a manually curated benchmark for evaluating the robustness and generalization of GUI grounding models under atypical in-the-wild visual observations. Additionally, when integrated into an agentic framework with a VLM planner, our approach improves success on end-to-end GUI tasks. Code, all data (training data and benchmark), and models will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.