acceptodds
Under review as a conference paper at ICLR 2027

Marionette: Explicit World State for Interactive World Model

Abstract

Interactive game world models typically generate video autoregressively in pixel or latent space. As a model repeatedly conditions on its own predictions, its inputs can drift away from the training distribution, causing errors to accumulate over long horizons. Correcting this drift is difficult when the dynamics remain implicit in visual latents. Our key insight is to identify the dynamics relevant to interaction and autoregress them in an explicit representation where geometric constraints can correct implausible predictions. In games with articulated characters, these dynamics take the form of 3D skeletal animation, which makes constraints on bone lengths, joint angles, and terrain contact directly expressible. We introduce Marionette, a world model that decouples explicit dynamics prediction from appearance generation. An action-conditioned model predicts character poses and trajectories; a deterministic graphics bridge applies state constraints and rasterizes the resulting animation into pose-control videos. A video diffusion model then renders the motion into the final frames. Experiments show that the predicted motion responds to action inputs and that state constraints reduce ground penetration while preserving skeletal structure during long rollouts, without changing the video model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.