acceptodds
Under review as a conference paper at ICLR 2027

AutoGUIWorld: Image Generators as Visual World Models for GUI Agent

Abstract

Scaling diverse training data for graphical user interface (GUI) agents is constrained by the environments available for interaction. Environment-based data engines typically construct platform- or task-specific interfaces and collect trajectories through execution, making broader coverage a recurring investment in development time and engineering effort. Yet screenshot-based agents learn from visible action consequences, motivating interaction fidelity as a data-generation target: preserving action-consistent visual feedback and task-state evolution without reconstructing the underlying software. The discrete, structured nature of GUI transitions makes image generators a natural candidate for implicit visual world models. We introduce AutoGUIWorld, an environment-free data engine that turns this premise into screenshot–action–screenshot trajectories. Structured world sampling combines platform, appearance, and interface-state constraints with Image2’s generative stochasticity to produce diverse visual seeds, from which tasks are generated. A meta planner then specifies atomic actions and their intended consequences; Image2 chain-edits the current screenshot into the next observation, while a grounding module annotates action targets. This factorization separates semantic planning from visual transition synthesis and expands data coverage without constructing or running the target applications.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.