MobileSkillBench: Automatically Constructing Memory-Dependent Mobile-Agent Tasks from Reusable Skills
Abstract
Personal mobile assistance requires agents to combine reusable operations with user information that determines which actions are appropriate. Evaluating this ability requires tasks in which personal records influence the workflow and its expected state. Existing work provides interactive environments, personalized tasks, and structured task synthesis, but constructing the dependencies between operations and user records remains a separate design problem. We present MobileSkillBench, a framework for constructing mobile-agent tasks from raw skill documents and de-identified user memories. A Skill–Memory Graph represents operation dependencies and the records that supply inputs, initialize application data, or constrain valid results. Joint Skill-Memory Search (JSMS) builds candidate subgraphs containing skill workflows and records from one user, then balances candidate quality against repeated use of the same graph resources. Auditable Task Compilation translates each selected subgraph into an instruction, initial state, expected state, and evaluator, with source evidence linking these components. Controlled recompilation checks whether removing a record changes task requirements or leaves essential information unresolved. The resulting MobileSkillBench comprises 281 tasks that evaluate LLM agents’ ability to use user memory to complete complex, multi-step tasks. Across eight current LLM agents, the best model achieves only 28.8% strict success and 67.9% Completion, with strict success falling to 10.9% on multi-app tasks. These results reveal substantial room for improvement in personalized mobile assistance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.