acceptodds
Under review as a conference paper at ICLR 2027

ConsisWorld: Benchmarking Spatial and Control Consistency in Interactive Video World Models

Abstract

Interactive video world models allow users to explore generated environments through movement and turning commands. Reliable interaction requires both consistent scene structure across viewpoints and the ability to anticipate subsequent action effects from prior experience. We introduce ConsisWorld to jointly evaluate spatial and action consistency through three tests: scene geometry preservation, single-action execution, and multi-action composition. The benchmark contains 140 initial views from 30 synthetic and real-world scenes, covering 36,400 videos generated by ten models. The spatial test constructs a fixed 3D reference from the initial image, depth, and camera parameters, and checks whether generated views remain consistent with the original scene geometry as the camera moves. The action tests execute each of six basic actions repeatedly, estimate its single-step displacement and rotation from the first four steps, and use these estimates to predict subsequent motion and the endpoint relations of independently generated composite routes. Each model is thus evaluated against its own measured action effects. We find that some models preserve scene geometry well yet exhibit substantial motion deviations during repeated execution of the same action. Composing different actions exacerbates these consistency challenges, causing actual endpoint relations to deviate significantly from the predictions of isolated actions. Moreover, some models are stable across repeated generations from the same initial view, but the displacement induced by an action changes substantially across initial views. These results show that good scene preservation does not imply stable continuous control. There remains substantial room for improvement in reducing motion deviations during sustained execution and action transitions, and improving control consistency across initial views. ConsisWorld evaluates scene preservation and action execution together, providing a benchmark for comparing model capabilities and assessing future improvements.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.