Seg2Click: Learning Pixel-Level Action Masks for GUI Visual Grounding
Abstract
Visual grounding is a central challenge for GUI agents, requiring an instruction to be mapped to a target within a crowded screenshot. Coordinate generation with single-point supervision reduces a spatially extended target to one position, despite the presence of many valid click locations. Region-based patch scoring alleviates this ambiguity, but its spatial decisions remain tied to visual-token footprints that may span the target, neighboring controls, and background. To address these limitations, we propose Seg2Click, which reformulates GUI grounding as executable-region prediction: identifying where a click can validly execute the given instruction. Using bounding-box supervision, Seg2Click predicts an executable-region mask and computes the final click point through a deterministic rule. This formulation captures the spatial extent of valid click locations and enables finer localization beyond the visual-token grid, making it well suited to small, densely arranged controls without requiring a verifier or zoom-in pass. Experiments across five backbone configurations from two model families demonstrate the effectiveness of Seg2Click. Seg2Click-9B achieves 92.06%, 94.10%, and 55.22% on ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro, respectively, while matched-backbone comparisons show improvements over both coordinate generation and patch-based grounding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.