acceptodds
Under review as a conference paper at ICLR 2027

AlayaWorld: Interactive Long-Horizon and Playable Video World Generation

Abstract

Building playable virtual worlds traditionally requires substantial manual effort in asset creation and interaction design. Video world models offer a generative alternative, but sustained interaction requires models to respond to user controls and perform diverse actions while maintaining long-horizon scene consistency and generation stability. We present AlayaWorld, an open video world model for long-horizon interactive and playable world generation, supporting camera-controlled navigation and text-driven behaviors such as spell casting during continuous generation. AlayaWorld generates video autoregressively in chunks, coordinating persistent spatial memory, temporal history, and motion-aware conditions to support scene recall across viewpoints and continuous evolution across chunks. To support navigation and interactive behaviors across diverse scenes, we curate and annotate six real-world and gameplay datasets under a unified protocol and further construct GenEvent, a synthetic video dataset that supplements this mixture with supervision for instruction-driven behaviors such as spell casting. We combine training on imperfect histories with autoregressive distillation to improve long-horizon generation stability and enable four-step sampling per chunk. Experiments on iWorld-Bench, WorldMark, and WBench demonstrate the effectiveness of AlayaWorld in world consistency, generation stability, and interactive control. We provide an open research framework with reference implementations and reproducible pipelines for interactive and playable video worlds.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.