acceptodds
Under review as a conference paper at ICLR 2027

A World Model Can Predict the Future Without Knowing Its Cause: Interventional Evaluation of Action-Conditioned World Models

Abstract

A world model exists to answer a planner's counterfactual question: if the action were different, how would the future differ? Most current evaluations score one generated future for its fidelity, plausibility or memory. Causal response, whether the predicted future changes correctly when the action changes, is instead a property of two futures. An elementary proposition makes this exact: no score that sees only a rollout and its context can measure causal response, and neither can any functional of the distribution of such rollouts. A reference-based score such as PSNR certifies it only when the prediction error is small compared with the true effect; under an equal-error assumption, this fails in most games for the models we test. We therefore measure the response directly, replaying paired interventions on an exact emulator under common random numbers. We study three world models of different design on 26 Atari games and a recurrent state-space model on six DeepMind Control tasks. On Alien the diffusion model's PSNR is within a decibel of its PSNR on Pong, and its consistency score is higher. Yet by delay 20 its response to a changed action is world-sized and has near-zero alignment with the world's. The alignment contrast holds under two intervention designs and in all three architectures. For the diffusion model it holds in all 96 analysis settings over 7 seeds and at a doubled horizon, and for the token model in at least 92 of 96. Causal difficulty transfers between every pair of architectures (rho=0.60 to 0.83, all surviving Holm correction), although absolute levels differ, and blindness is a property of a model on a game under a context distribution. Two instruments meant to characterise the failure did not survive their own checks. A latent probe reports an encoded consequence that the same probe does not find once the action pathway is held fixed, and an automated object-level readout is not validated by a blinded audit with three annotators. We therefore report both as audited instruments, not as findings. Alignment's association with planning usefulness remains inconclusive. On Alien, where every model is blind, no model shows a detectable gain over a random policy, yet elsewhere models with low alignment under the diagnostic protocol can still improve returns.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.