EASE: Efficient Mobile Agents with Self-Maintained State and Executor-Calibrated Planning for Long-Horizon Tasks
Abstract
Graphical user interface (GUI) agents have achieved strong performance by integrating perception, reasoning, and action generation within a single model, yet still struggle with long-horizon and complex tasks. Hierarchical systems address these challenges through specialized modules for planning, memory, and execution, but incur substantial inference overhead. This creates a fundamental trade-off between efficient end-to-end execution and higher-overhead hierarchical assistance. We address this trade-off with EASE-Agent, a mobile GUI agent that preserves lightweight execution while leveraging sparse strong-model assistance to improve long-horizon task performance. To realize this design, we introduce two complementary mechanisms. First, we equip the executor EASE-UI with Self-Maintained State and introduce Token-level Advantage Policy Optimization (TAPO) to reduce cross-objective interference during joint training. Second, we introduce executor-calibrated Skill Memory for task decomposition, enabling the planner to adapt subgoal granularity to the executor's capabilities. Extensive experiments demonstrate the effectiveness and efficiency of our approach. On GUI-Odyssey, EASE-UI achieves the best performance among the compared methods, with TAPO outperforming GRPO and GDPO. On MemGUI-Bench, EASE-Agent achieves 31.3% pass@3, while stronger configurations further improve performance: EASE-Agent-Pro reaches 50.0%, and EASE-Agent-Max reaches 75.0%, compared with 49.2% for Agent-S2. EASE-Agent-Max incurs only 11.1% of Agent-S2's API cost per attempt.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.