DreamUI: Scaling Interactive Mobile Agent Environments with Generative UI World Model
Abstract
Training mobile agents for diverse, long-horizon tasks requires scalable interactive environments. However, real apps require authentication and permissions and are difficult to reset, while mock apps rely on predefined data and transition logic that limit task coverage and interaction diversity. In this paper, we present DreamUI, a scalable reinforcement learning framework that leverages a Generative UI World Model as the training environment. Our core idea is to generate diverse environments while preserving task consistency. The UI world model preserves task-essential facts while introducing plausible variations in unspecified content, exposing agents to richer training contexts. It also generates task-relevant states conditioned on instructions, reducing reliance on predefined app databases when expanding training tasks. Given a task instruction, the current observation, and an agent action, the UI world model generates HTML that renders into the next UI observation, forming a closed interaction loop. To address imperfect environment predictions and sparse task-level feedback, we introduce Environment-Aware Trajectory Masking to exclude corrupted rollouts and a Rubrics-Based Reward System that provides graded rewards through fine-grained task criteria. We train the UI world model through supervised fine-tuning followed by step-wise reinforcement learning. Our DreamUI-World-8B achieves average Instruction Accuracy and Visual Similarity scores of 82.9% and 75.6%, respectively. As for mobile-agent training, RL in the simulated environment raises our DreamUI-4B's success rates to 73.3% on AndroidWorld and 36.2% on MobileWorld. These results demonstrate that our UI world models can serve as effective training environments for mobile-agent RL. We will release our code, data, and models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.