TAP-Agent: From One-Way Handoffs to Two-Sided Preparation for Device–Cloud GUI Agents
Abstract
Capable graphical user interface (GUI) agents increasingly rely on cloud-hosted multimodal large language models for high-level planning and lightweight on-device models for visual grounding and execution. However, existing device–cloud frameworks treat each cloud response as a rigid handoff from planning to execution. This serialization forces the device to ground targets only after a plan arrives while leaving its computation underutilized during cloud reasoning, resulting in inaccurate execution and increased interaction latency. To address these limitations, we propose **TAP-Agent** (**T**wo-sided **A**nticipatory **P**reparation), which prepares on-device computation on both sides of each cloud response. After a response arrives, plan-conditioned visual experts provide heterogeneous visual evidence, which is compiled into action-ready representations that associate semantic elements with screen geometry. While awaiting the next response, a compact experience buffer of verified interactions guides the device to anticipate likely operations and prefetch supporting visual evidence. An alignment mechanism reuses this computation when it is consistent with the returned plan, while UI-changing actions remain governed by confirmed instructions. Together, these mechanisms coordinate computation across time, improving execution reliability while shifting reusable device inference off the critical path. Experiments on AndroidWorld and OSWorld with multiple on-device executors show that TAP-Agent improves task success by **5.7%** and reduces end-to-end interaction latency by **10.4%** over state-of-the-art device–cloud baselines. The code will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.