Enforcing Causal Routing in Embodied World Models
Abstract
Embodied world models predict how a robot's environment responds to commanded actions, enabling policy evaluation and planning with varied actions. We show that current models produce such predictions partly by peeking at the future. Frames generated at early timesteps can be corrupted by subsequent actions and the goal, a phenomenon we term **causal leakage**. Existing benchmarks typically replay recorded actions, which match the ground-truth video, so leakage does not lower scores and goes undetected. We introduce Causeway, a world model that routes information along causal paths while preserving visual quality. United Causal Attention aligns the attention structure with this causal order while retaining semantic guidance from language; Biomimetic Sampling Memory sustains prediction fidelity over episode-length rollouts by anchoring the initial observation alongside recent context. To make causal leakage measurable, we introduce the Causal Leakage Probe, to our knowledge the first black-box test of whether subsequent actions affect earlier frames or the goal affects the video directly. Causeway's future-action leakage is 63% below the best baseline's, with no measurable goal leakage under the tested interventions. It leads on six of twelve WorldArena 2.0 metrics and reduces FVD on full-episode DROID rollouts from 113.6 to 8.9 over the same backbone. Code and weights will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.