acceptodds
Under review as a conference paper at ICLR 2027

Wonder: Video World Model Done Better

Abstract

We present Wonder, a general-purpose video world model for real-time, camera-controllable world exploration. Given a single image or a conditioning video, Wonder constructs a playable world in which users navigate by moving the camera, discovering unseen regions and revisiting observed ones over a long horizon. This requires a system-level co-design of control, memory, and training. We introduce a camera conditioning scheme based on a dense pixel-space coordinate field whose renderings supply spatially aligned motion and orientation cues, so the model reads camera motion directly as visual evidence. For fast retrieval over a growing context, we propose a sparse full-fidelity memory mechanism that attends to a small set of relevant tokens at inference, independent of context length. We further rectify the self-forcing distillation pipeline with a timestep-wise mixture-of-students and a camera-aware adversarial regularizer, improving control adherence while retaining the teacher's generation diversity and stability. Together, these components let Wonder synthesize diverse, minute-scale videos at 16 FPS with coherent geometry, appearance, and dynamics across long rollouts. Beyond image-to-video, Wonder natively supports video-conditioned generation, re-shooting existing dynamic scenes in real time.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.