A Discrete World Exploration Model with Superior Action Alignment
Abstract
As a fundamental world-modeling task, world exploration requires accurate action alignment and stable long-horizon generation. Existing approaches are predominantly diffusion-based but still exhibit limitations, particularly in accurate action alignment. Moving beyond diffusion, we introduce DiscreteWorld, a discrete generative model that produces stable 10-second, 480p, camera-controllable open-world videos with competitive action controllability. A key design of DiscreteWorld is to quantize both camera actions and visual observations into a unified discrete token space for joint reasoning and generation, yielding two essential advantages. First, unified action-vision modeling in the same token space enables precise semantic alignment between the two modalities, leading to accurate and reliable action following. Second, viewing discretization as a form of refinement and compression, we conduct the entire world modeling—including semantic understanding and video generation—in a discrete token space to facilitate the learning of higher‑quality latent compressed representations, which is crucial for world models. Additionally, we build our DiscreteWorld upon GRNs—a text‑to‑video discrete generative model—and thereby naturally inherit its iterative visual refinement capability for high‑quality video synthesis. Extensive experiments on real-world and game scenarios show that DiscreteWorld achieves more accurate action following and more stable long-horizon generation than prior diffusion-based world models. All models and code will be open-sourced to foster further research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.