acceptodds
Under review as a conference paper at ICLR 2027

Injecting Causal Structure into Video World Models

Abstract

World models should answer "what-if" questions about interventions on an environment. We introduce Causal Markov Supervision (CMS), a post-training procedure that injects a user-specified causal DAG into a video world model to improve intervention simulation. Each training episode is labeled with an event trace of DAG-variable outcomes and their timing, and a pre-trained video world model learns to predict traces along with predicts frames and actions after post-training. The DAG supervises temporal ordering and constrains attention so trace predictions satisfy its Markov property. In a realistic, AAA-quality Unity environment, confounding by player choice causes standard world-model training to fail on the causal effect of a "trapped chest” mechanic on level completion. CMS accurately simulates identifiable intervention outcomes without degrading pretrained visual fidelity. With intervention labels, it also exploits intervention-induced independences to simulate beyond the observational training distribution. Finally, we show that a vision–language model can automatically label training episode with DAG event traces.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.