AutoFocus: Uncertainty-Aware Active Visual Search for GUI Grounding
Abstract
Vision-Language Models (VLMs) have enabled autonomous GUI agents that translate natural language instructions into executable screen coordinates. However, grounding performance degrades in high-resolution interfaces, where dense layouts and small interactive elements expose a resolution gap between modern displays and model input constraints. Recent methods have explored zoom-in refinement, confidence-based selection, and reward-driven optimization. But most of them either rely on manually designed crop policies, consensus over sampled predictions, or additional reinforcement learning objectives. In contrast, we study whether the model’s own autoregressive coordinate distribution can serve as an intrinsic signal for active visual search at test time. We propose AutoFocus, a training-free uncertainty-aware refinement framework that converts token-level coordinate perplexity into an anisotropic spatial probability field. This field guides global-local region proposal generation, while an error-triggered verifier selectively activates refinement only when the initial prediction is uncertain. AutoFocus therefore improves high-resolution GUI grounding without requiring additional annotations, reward modeling, or parameter updates. Extensive experiments on ScreenSpot-Pro, ScreenSpot-V2 and OSWorld-G demonstrate consistent improvements across both general-purpose and GUI-specialized VLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.