acceptodds
Under review as a conference paper at ICLR 2027

MemGUI-RL: Reinforcement Learning for Proactive Context Management in Long-Horizon Mobile GUI Agents

Abstract

Long-horizon mobile GUI agents must decide not only which button to tap but also what to remember: which history to compress, which screen facts to keep, and when. **MemGUI-Agent** casts these decisions as first-class actions through the text-as-ion (**CONACT**) interface, but its supervised policy only imitates demonstrations and never compares alternative memory decisions sampled from itself. We study reinforcement-learning post-training of such policy-managed memory and find that group relative policy optimization (**GRPO**), applied naively, fails in two ways. ***First***, aggregating heterogeneous verifier rewards (format, action type, parameters, folding range) before group normalization lets the component with the largest within-group spread dictate the update, regardless of the intended weights. ***Second***, span-level history folds, the operation that actually compresses context, are only 22.7% of annotated folds, and the majority, step folds, are almost always already correct, so under natural sampling most prompt groups receive identical rewards and the policy collapses to step-only folding (deep-fold rate 17.7% → 0.2%). ***(i)*** We introduce olding-ware eward-Decoupled olicy ptimization (**FARPO**), which standardizes each verifier component within its group before weighting and controls span exposure through a span-to-step ratio ; both mechanisms are characterized analytically. ***(ii)*** The resulting **MemGUI-8B-RL** reaches 48.4% Pass@3 and 39.0% information retention on MemGUI-Bench and 19.7% success on the out-of-distribution MobileWorld, more than doubling the backbone model on every metric, with the gain concentrated on memory-intensive tasks; it is the best open-data 8B model on both benchmarks. ***(iii)*** A seven-ratio sweep shows folding exposure to be a controllable hyper-parameter with a bimodal footprint, and a stepwise ablation attributes the gain to each mechanism. ***An anonymized release of the code, data, trained model, and evaluation logs is publicly available at [https://memgui-rl-anonymous.github.io/](https://memgui-rl-anonymous.github.io/).***

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.