acceptodds
Under review as a conference paper at ICLR 2027

UI-Mem: Stratified Memory-Guided Online Reinforcement Learning for Mobile GUI Agents

Abstract

Online Reinforcement Learning (RL) offers a promising paradigm for enhancing mobile GUI agents through direct environment interaction. However, its effectiveness is severely hindered by long horizons and sparse rewards: under group-based methods such as GRPO, sampled rollout groups often collapse into uniform failure, yielding near-zero advantage and little relative learning signal. We propose UI-Mem, a framework that repurposes past experience as a training-time scaffold for online RL. Instead of merely replaying successful trajectories or scoring rollouts after collection, UI-Mem uses stratified memory-guided group sampling: within each GRPO group, retrieved experience is injected at different strengths to produce trajectories with full, partial, or no guidance. These contrasted rollouts create informative reward differences and train the unguided policy to internalize the external experience. UI-Mem instantiates this scaffold with a hierarchical experience memory, where workflows, subtask skills, and failure patterns are stored as parameterized templates and updated online from newly collected successes and failures. Experimental results show that UI-Mem substantially outperforms vanilla RL, dense-reward training, and static experience reuse across model scales, with strong generalization to unseen applications.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.