AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
Abstract
Scaling diverse training data for graphical user interface (GUI) agents is constrained by the environments available for interaction. Environment-based data engines typically construct platform- or task-specific interfaces and collect trajectories through execution, making broader coverage a recurring investment in development time and engineering effort. Yet screenshot-based agents learn from visible action consequences, motivating interaction fidelity as a data-generation target: preserving action-consistent visual feedback and task-state evolution without reconstructing the underlying software. The discrete, structured nature of GUI transitions makes image generators a natural candidate for implicit visual world models. We introduce AutoGUIWorld, an environment-free data engine that turns this premise into screenshot–action–screenshot trajectories. Structured world sampling combines platform, appearance, and interface-state constraints with Image2’s generative stochasticity to produce diverse visual seeds, from which tasks are generated. A meta planner then specifies atomic actions and their intended consequences; Image2 chain-edits the current screenshot into the next observation, while a grounding module annotates action targets. This factorization separates semantic planning from visual transition synthesis and expands data coverage without constructing or running the target applications.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.