acceptodds
Under review as a conference paper at ICLR 2027

Benchmarking General Mobile Assistants in Challenging Real-World Scenarios

Abstract

Graphical user interfaces provide a challenging setting for evaluating agents that must perceive, reason, and act over interaction trajectories. Existing mobile benchmarks have established reproducible Android evaluation environments, but there remains room to broaden application coverage and capture more demanding user workflows. We present GMA, a benchmark for evaluating general mobile assistants in challenging real-world scenarios across seven custom applications adapted from open-source projects and 300 tasks spanning four progressive difficulty tiers, from atomic interactions to realistic long-horizon workflows. Evaluation of nine frontier models under a shared agent setup reveals a sharp decline in task completion as execution demands increase, with agents increasingly making partial progress without completing the full workflow. We further find that preserving interaction history substantially improves performance, especially on demanding tasks, while retaining increasingly more raw context eventually yields diminishing returns. Motivated by these results, we use GMA as a controlled testbed for mobile harness design, covering context management, within-task memory, planning, and verification. Our experiments show that harness effectiveness is highly conditional. The effect of a given design varies with task demands, alternative implementations within the same module can yield substantially different outcomes, composition can diminish or even reverse marginal gains, and the same design can affect foundation models differently. Overall, GMA provides a reproducible testbed for mobile agent evaluation and a controlled setting for understanding how foundation models and their surrounding harness jointly shape reliable execution in complex mobile workflows.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.