Beyond Decoding Acceleration: Fast Plan-Conditioned Grounding for GUI Agents
Abstract
Multimodal GUI agents can automate web, desktop, and mobile interfaces from natural language instructions and screenshots, but their practical use is limited by high per-step latency. While existing acceleration methods mainly reduce subsequent-token decoding cost, GUI grounding typically produces only short actions and coordinates. Our latency analysis shows that the dominant bottleneck is instead the first-token latency caused by repeated visual encoding, multimodal fusion, and prefilling in large vision-language models. To address this, we propose PACE, a Plan-conditioned Action and Coordinate Executor, within a two-speed GUI agent framework. A main model generates concise next-step plans, and PACE executes simple and safe steps by directly predicting the action type and target location from the plan and screenshot. PACE is a lightweight 0.2B dual-stream model with cross-modal fusion and a grid-based grounding head. Across five GUI grounding benchmarks, PACE achieves 64.0% average accuracy with 46.59 ms latency. On OSWorld, integrating PACE reduces main-model calls and total execution time while largely preserving task performance. These results show that low-latency GUI agents require not only faster decoding, but also plan-conditioned delegation that avoids unnecessary heavy multimodal computation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.