MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management
Abstract
MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet remain unreliable on long-horizon tasks that require retaining intermediate facts across many steps and app transitions. We attribute this limitation to ReAct-style prompting, which passively accumulates per-step records, leading to prompt explosion and dilution of critical cross-app facts. To address this, ***(i)*** we introduce **MemGUI-Agent**, an end-to-end long-horizon mobile GUI agent with proactive context management. MemGUI-Agent is built on text-as-ion (**CONACT**), which casts context management as first-class actions emitted by the same policy that selects UI actions. Instead of passively appending history, CONACT actively maintains three structured context fields: folded action history, folded UI state, and recent step record, enabling the agent to preserve critical UI facts while keeping context compact. To make proactive context management learnable across model scales, ***(ii)*** we construct **MemGUI-3K**, a 2,956-trajectory dataset with full CONACT annotations, enabling supervised training and analysis of proactive context management. ***(iii)*** Training an 8B model on MemGUI-3K produces **MemGUI-8B-SFT**, an 8B MemGUI-Agent that achieves the best open-data 8B performance on MemGUI-Bench and demonstrates strong out-of-distribution performance on MobileWorld. ***An anonymized release of the code, data, and trained models is publicly available at [https://memgui-agent-anonymous.github.io/](https://memgui-agent-anonymous.github.io/).***
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.