SeerWorld: Towards Out-of-Sight World Evolution in Generative Video Models
Abstract
Interactive video world models offer a promising route toward world simulation by enabling camera-controlled exploration of evolving visual environments, while aiming to maintain consistency across changing viewpoints. Faithful world simulation, however, requires more than preserving what has previously been observed: entities should remain participants in the world's dynamics, interacting with events and changing in response even outside the observer's field of view. Current models struggle throughout this process: a mid-rollout prompt switch may fail to instantiate the new event; when the camera later revisits an earlier region, the subject and background may drift or disappear; and even when the subject is preserved, it often remains behaviorally static rather than reacting to or interacting with the event, making the resulting world evolution behaviorally and causally implausible. We formalize this challenge as Out-of-Sight Reactive Evolution and introduce the SeerWorld-Data along with SeerWorld-Bench. As a baseline reference solution, we develop SeerWorld, a segment-autoregressive video model trained on our proposed data, together with a reasoning-augmented generation framework, which autonomously reasons about how the off-screen world evolves and renders the resulting dynamics. Evaluation results show that our approach outperforms recent state-of-the-art models in event-conditioned interaction rendering while preserving revisit consistency, establishing a reference point for autonomous world evolution beyond the current field of view. The model and data will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.