Seeing the Forest and the Trees: Multi-Scale Voting Improves Test-Time GUI Grounding
Abstract
GUI grounding, the localization of screen elements based on natural language instructions, is essential for autonomous GUI agents. Supervised training achieves strong accuracy but requires massive labeled data, making annotation expensive and limiting generalization to new environments. Recent work addresses this challenge by exploring test-time grounding, which exploits unlabeled data available during inference via iterative cropping and spatial voting. In this paper, we identify scale variation as an important yet overlooked factor in test-time GUI grounding and propose a simple Multiscale Spatial Voting (MSV) method. Specifically, MSV scales the full image up and down to multiple resolutions, generates multiple predictions at each scale, and aggregates these predictions via spatial voting to produce the final output. We integrate MSV into various test-time grounding approaches, including three training-free methods and two training-based methods. Extensive experiments on four datasets with various backbones confirm that MSV consistently improves grounding performance. Despite its simplicity, we believe this work provides a promising path toward label-efficient GUI agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.