acceptodds
Under review as a conference paper at ICLR 2027

ZoomSD: Learning Where to Zoom in Self-Distilled GUI Agents

Abstract

GUI grounding—locating the target element of a natural-language instruction on a screenshot—is a core capability of GUI agents. While training algorithms such as on-policy self-distillation (OPSD) have greatly improved grounding, models still struggle on high-resolution screenshots with small icons. Zoom-in inference mitigates this problem by cropping around a first-pass prediction and re-encoding the enlarged crop, but the closer view comes at a price: the model must first decode a coarse coordinate in an extra round of inference, roughly doubling the latency. This raises a natural question: is explicit coordinate generation necessary for deciding where to zoom? We demonstrate that a few attention heads involved in coordinate decoding already provide a useful coarse localization signal. Building on this observation, we propose ZoomSD, with two components. ILR Crop Selection applies bounded low-rank corrections (ILR) to the selected heads' attention logits and selects a fixed-size crop from the refined attention map. Second, Joint Training with OPSD and Containment Ranking couples this selector with OPSD through a containment ranking loss, with gradients isolated from the grounding backbone. At inference time, one full-image prefill selects the crop without autoregressive coordinate decoding; the enlarged crop is then used for final click generation. On Qwen3-VL-4B and Qwen3-VL-8B, ZoomSD substantially improves over the vanilla OPSD baseline on ScreenSpot-Pro, UI-Vision, and OSWorld-G, matching or even exceeding two-stage zoom-in accuracy while reducing its inference latency by 12.4–32.5%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.