acceptodds
Under review as a conference paper at ICLR 2027

From Benchmarks to the Wild: Practical Design Choices for Deployable Mobile GUI Agents

Abstract

Recent years have witnessed rapid advances in large multimodal model (LMM)-based mobile GUI agents that have achieved strong performance on related benchmarks such as ScreenSpot and AndroidWorld. However, it remains challenging to deploy such benchmark-capable agents in real-world scenarios where they must operate under dynamic UI states, diverse user instructions, third-party application variations, and strict latency budgets that existing benchmarks often underrepresent. To investigate how to bridge this gap, we present an empirical study of practical mobile GUI agent design through deployment-oriented experiments across Android and HarmonyOS, covering 50+ mainstream applications. Our analysis focuses on two important dimensions: data construction strategies that make agents robust beyond benchmark-like settings, and latency-aware designs that keep accurate agents responsive as well. For robustness, to handle the wide variation of user inputs and environment states in real use, we propose three complementary data construction schemes that reflect diverse, realistic workloads and transfer GUI agents from fixed benchmarks to dynamic scenarios, improving real-world task success rate by 10.2 percentage points. For efficiency, we find that heavy reasoning or multi-agent workflows emphasized by existing agent paradigms can introduce unacceptable latency; by streamlining agent input/output and workflows, we identify the Pareto-optimal inference-time designs for balancing task accuracy and latency. Together, these findings provide practical guidelines for building mobile GUI agents that are robust, efficient, and deployable in real-world environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.