SCOPE: Spatial Confidence and Offloading with Per-Element Guarantees for Device-Cloud GUI Grounding
Abstract
GUI grounding turns a natural language instruction into a coordinate on the screen. Since cloud inference is costly, practical deployments run lightweight edge agents on the device. The edge agent closes the capability gap by offloading the tasks it would get wrong, which calls for uncertainty estimation and a risk-controlled offloading decision. Existing methods typically estimate the uncertainty from logit statistics or repeated sampling, and offload by thresholding the score. However, these methods measure hesitation implicitly without quantifying the deviation of the prediction from the target element, which decides the correctness of grounding. We instead expand the logits of the grounding coordinate digits into a click distribution over the screen, and observe that the hesitation separating correct from wrong predictions lies at the high-order digits that decide the element rather than at the low-order digits, and that wrong groundings often escape to the neighborhood of the target. Building on this, we propose SCOPE, a framework for Spatial Confidence and Offloading with Per-Element Guarantees. SCOPE first reads the separation of hesitation across coordinate scales through an entropy gate, cross-checks the prediction by re-grounding on a local view, and calibrates the combined spatial confidence into a conformal click radius in units of UI elements. The edge offloads with a per-element guarantee whenever the radius exceeds one element. Experiments demonstrate that the spatial confidence recovers on average 7.7 percentage points more of the device-cloud capability gap than the strongest baseline, and the calibrated radius offloads 58% of inputs at 90% nominal coverage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.