Mobile Skills Bench: Benchmarking Long-Horizon Interactive Mobile Agents with Multimodal Skills under Hybrid Execution
Abstract
Mobile agents are increasingly expected to carry out everyday workflows that span multiple applications, changing interfaces, and evolving user requests. However, such long-horizon tasks remain challenging for current agents. Agent skills offer a promising remedy, because they package reusable procedural knowledge, and multimodal skills further add visual grounding through reference images. Yet their benefit remains unclear, since no benchmark compares skill designs under controlled conditions on long, interactive mobile tasks with verifiable outcomes and hybrid execution. To address this gap, we introduce Mobile Skills Bench (MSB), a controlled testbed for evaluating agent skills, which comprises 145 long-horizon tasks across 24 simulated applications with a verifiable backend, provides a user simulator, and supports both Pure GUI and Hybrid (GUI, MCP, and CLI) execution. MSB further includes a multimodal skill library of 73 procedures with 166 reference images, organized by function, by function–app pair, and by app. We then evaluate nine models in three agent harnesses. Without skills, the nine models achieve an average task success of only 28.9% under Hybrid and 12.5% under Pure GUI execution. The models often fail to ask for missing information, claim completion prematurely, or stall without recovery. With textual skills, the two averages rise to 34.2% and 15.0%, although the gain varies with the model and execution mode. However, adding the tested static reference images yields no consistent further gain. Finally, the three library organizations reach similar task success, but organizing the library by app uses 31% fewer input tokens on average than organizing it by function or by function–app pair, which suggests a simple principle for designing skill libraries in mobile scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.