UI-Mem: Stratified Memory-Guided Online Reinforcement Learning for Mobile GUI Agents
Abstract
Online Reinforcement Learning (RL) offers a promising paradigm for enhancing mobile GUI agents through direct environment interaction. However, its effectiveness is severely hindered by long horizons and sparse rewards: under group-based methods such as GRPO, sampled rollout groups often collapse into uniform failure, yielding near-zero advantage and little relative learning signal. We propose UI-Mem, a framework that repurposes past experience as a training-time scaffold for online RL. Instead of merely replaying successful trajectories or scoring rollouts after collection, UI-Mem uses stratified memory-guided group sampling: within each GRPO group, retrieved experience is injected at different strengths to produce trajectories with full, partial, or no guidance. These contrasted rollouts create informative reward differences and train the unguided policy to internalize the external experience. UI-Mem instantiates this scaffold with a hierarchical experience memory, where workflows, subtask skills, and failure patterns are stored as parameterized templates and updated online from newly collected successes and failures. Experimental results show that UI-Mem substantially outperforms vanilla RL, dense-reward training, and static experience reuse across model scales, with strong generalization to unseen applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.