Empowering World Models with Deconfounded Dynamics via Multi-Level Causal Alignment
Abstract
Implicit world models predict future visual representations without reconstructing images, supporting tasks such as embodied planning. However, encoding object and background information into a shared feature space can entangle task-relevant states with incidental visual context. This may bias future-state predictions and propagate errors over successive predictions. To address this problem, this paper presents a Deconfounded Causally Aligned Latent Model (DeCALM), which empowers world models with deconfounded dynamics by reducing background dependence in latent encoding. Specifically, DeCALM comprises two main modules: Background-Deconfounded Latent Encoding (BLE) and Multi-Level Causal Alignment (MCA). BLE uses object and background representations separately, approximates backdoor adjustment with a background dictionary, and constrains predictions under counterfactual background changes. MCA takes the resulting features and adaptively combines absolute-state and incremental predictions to improve future-state prediction. Experiments on three benchmark datasets show that DeCALM achieves the highest success rate among the compared methods, including a 14.58-percentage-point gain over the strongest baseline on Drawer Open. Further analyses suggest that DeCALM mitigates the ill-posedness of latent representations under visual disturbances, yielding more task-relevant features and more accurate state-change predictions at later horizons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.