Lazy MCFV: Observation-Aware Selection with Exact Query Pruning for Multi-view GUI Grounding
Abstract
GUI grounding connects natural-language instructions to actionable locations in interface screenshots. Multi-view observation and magnification improve the perception of small interface elements, yet different views may produce conflicting predictions. Existing aggregation methods treat these predictions as isolated coordinates and discard the observation context from which they originate. We argue that the generating view is an essential component of a grounding prediction. Minimal Common Field-of-View (MCFV) is a training-free, observation-aware selection rule that preserves this provenance by pairing each candidate with its generating view and selecting the finest view whose observation scope contains all candidate locations. It automatically retains detailed local observations when they are sufficient and broader contextual views when ambiguity requires additional coverage. The same containment structure enables Lazy MCFV to achieve exact query pruning without changing the selector's output. We establish exact output preservation and show that the required observations are determined by the selector itself rather than heuristic confidence estimates. Experiments on ScreenSpot-Pro demonstrate improved selection from shared observations and output-preserving query reduction without additional training or verification.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.