MobiGraph: Towards a Practical Evaluation Paradigm for Mobile Agents via Trajectory-Fused State Graphs
Abstract
Mobile agents can autonomously complete user-assigned tasks through GUI interactions. However, existing mainstream evaluation benchmarks, such as AndroidWorld, operate by connecting to a system-level Android emulator and provide evaluation signals based on the state of system resources. In real-world mobile agent scenarios, many third-party applications do not expose system-level APIs to determine whether a task has succeeded, leading to a mismatch between benchmarks and real-world usage. To address these issues, we propose MobiGraph, an evaluation framework that supports task construction for arbitrary third-party apps. Using an efficient graph-construction algorithm based on multi-trajectory fusion, MobiGraph can effectively compress the state space, support dynamic interaction, and better align with real-world third-party application scenarios. MobiGraph covers 20 widely used third-party applications and comprises 240 diverse real-world tasks, with enriched evaluation metrics. Compared with AndroidWorld, MobiGraph 's evaluation results show higher alignment with human assessments and can guide the evolution of future GUI-based models under real workloads.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.