Edge Symphony: A Harness-Aware Benchmark for Hybrid Mobile GUI Agents
Abstract
Mobile GUI agents that run entirely in the cloud raise privacy concerns and incur substantial API cost. Hybrid agents instead distribute inference between on-device and remote models through an agent harness, which decides where computation runs, what information is exchanged, and when control transfers between models. This raises a central question: what capability does a harness realize in hybrid mobile GUI execution, and at what deployment cost? We introduce Edge Symphony, a harness-aware benchmark that evaluates hybrid mobile GUI agents on physical smartphones with native on-device inference. Edge Symphony comprises 258 reproducible tasks across 23 applications, organized into six capability categories and extensible through an automated task curation workflow, and jointly measures Smartness and Cost. We also provide two reference harnesses, PocketAct (fully local) and RelayAct (hybrid). Experiments with Edge Symphony reveal three findings on hybrid agents invisible to model-centric evaluation: efficiency in converting model calls into actions bounds realized capability; harness design relocates cost rather than eliminating it; and model gains does not guarantee a better agent, which is also determined by the harness. These findings provide valuable insights into hybrid mobile GUI agents design.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.