Tracing How Robot World Models Hallucinate Objects
Abstract
Action-conditioned video diffusion models now render robot motion convincingly, yet the objects the robot manipulates are often left behind, duplicated or lost. Standard video metrics miss these failures and even reward them, since a clip that repeats its first frame beats the generated video on temporal-difference error in 98% of 540 cases. Rather than adding data or conditioning, we present, to our knowledge, the first mechanistic account of how a robot video world model carries objects, using causal interventions in frozen diffusion transformers (DiTs). The control renders only the arm and steers it through attention routing, since restoring only the queries and keys (Q/K) of the correct control removes 84–88% of the error a corrupted control adds. The object appears only in the first image, so the model must copy it forward. We find that rotary position embedding (RoPE) turns the first image's temporal index into an address for this copy, and shifting that index alone makes the video return to its starting scene inside the predicted frame windows, in 10 of 10 held-out episodes and in two public video models. The copy is already visible in the model's first denoising step and is stored in the partially denoised latent, which redraws the object at its old place after the hand has moved it. Guided by this mechanism, we paint the object into the control, which makes it follow the hand in 6 of 16 new failures, and in our 14B model re-noising a leftover copy's region early removes the copy while keeping the carry. These findings trace a basic limitation of current robot world models to the fact that they are told where objects are only in the first image, and point future models toward conditioning on object state, acting from the first denoising steps and evaluating whether objects go where the actions send them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.