acceptodds
Under review as a conference paper at ICLR 2027

MoTWorld: Geometry-Grounded World Modeling with Joint Video-Voxel Generation

Abstract

Interactive world models predict future observations from actions of an agent, but maintaining spatial consistency over long rollouts remains challenging. Explicit 3D states provide a persistent representation of the scene, yet cascaded approaches typically predict geometry first and then use it to condition video generation, limiting information exchange between the two modalities. We propose MoTWorld, a world model that jointly generates video frames and semantic voxel grids within a Mixture-of-Transformers framework. Specifically, we represent images, voxels, and camera parameters as separate streams with modality-specific weights and shared attention. By changing which streams are observed or generated, we use the same model for image-to-voxel reconstruction, voxel-to-image rendering, and action-conditioned rollout. We further identify that shared attention alone does not explicitly establish correspondence between image and voxel tokens, which occupy different coordinate systems. We propose a novel rotary positional encoding that uses camera geometry and predicted depth to represent both modalities in a common 3D coordinate system. We account for voxel extent and depth uncertainty in the encoding to guide cross-modal attention. Experiments on two Minecraft-style test sets with exact 3D ground truth shows that MoTWorld significantly improves single-image voxel reconstruction and video quality over 600-frame rollouts compared with state-of-the-art methods, while supporting reconstruction, rendering, and generation within a unified architecture.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.