MobiCLI: Benchmarking CLI and Hybrid GUI-CLI Interaction for Mobile Agents
Abstract
Mobile agents have largely been developed and evaluated through GUI-only interaction, overlooking command-line interfaces (CLIs) that can improve execution efficiency and support operations beyond GUI actions. To bridge this gap, we introduce MobiCLI, a benchmark for evaluating CLI and hybrid GUI–CLI interaction on mobile devices. We identify five representative mobile scenarios where CLI capabilities are particularly useful: Batch Operations, File Management, Computational Tasks, Data Processing, and System Operations, and construct 122 tasks across 18 applications. We evaluate strong general-purpose models and representative phone-use agents on MobiCLI, with the best model achieving only 55.7% success. Our experiments reveal three key findings. First, CLI access substantially improves both success and efficiency for strong general-purpose models, exposing the limitations of GUI-only interaction. Second, existing phone-use agents generally fail to effectively exploit CLI capabilities. Third, hybrid tasks are substantially harder than CLI tasks, highlighting the challenge of GUI–CLI coordination. To improve hybrid execution, we further introduce MobiMux, a mobile agent framework that combines task verification with continual experience learning. MobiMux improves success rates by up to 16.4 percentage points on general-purpose models and achieves the best overall performance on MobiCLI.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.