WorldEngine: Reverse Engineering Visual Worlds with Executable Coding Agents
Abstract
Adapting AI models to novel visual domains typically requires humans to hand-engineer simulators, training data, and evaluators, whereas learned neural world models remain difficult to inspect, verify, or control. We introduce WorldEngine, a framework in which a coding agent turns a minimal task description and sparse visual examples into an executable world hypothesis: an engine that generates task instances, renders solution trajectories, and validates candidate solutions, with configurable difficulty and appearance. In Part I, evaluations across four visual reasoning and planning domains expose an executability-fidelity gap: engines that run without error still encode wrong rules, leak answers, or collapse controls into fixed templates. Our structured construction workflow improves rule fidelity and validity, and explicit diversity objectives activate controllable variation. In Part II, we use a synthesized engine to adapt a video-based visual planner in two regimes. Used open-loop, the engine supplies verified trajectories for fine-tuning, raising the solving rate on an independent simulator's puzzles from 2% to 24%. Used closed-loop, its validator feeds per-difficulty success back into sampling from a fixed pool, improving mean hard-level validation solving rate by up to 12.0 points over a static mix. Together, these results establish executable world reverse engineering as an inspectable, controllable engine for adapting visual planners.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.