Causal Nuisance Interventions for Distractor-Robust World Models
Abstract
World models trained from pixels can remain sensitive to camera motion and changing backgrounds, even when trained without pixel reconstruction. We introduce Causal Nuisance Interventions (CaNI), a causal formulation and replay-training method using a specified nuisance structural causal model (SCM) to guide observation-proxy design before training. We instantiate CaNI on a decoder-free Dreamer backbone by applying one randomly sampled padded translation consistently across each replay sequence, while retaining the architecture and learning objectives. We prove that uniform posterior stability is sufficient to bound changes in a fixed recurrent policy's return under observation interventions that preserve the controlled process. Across five DeepMind Control tasks with driving-video backgrounds and camera jitter, CaNI achieves higher mean return than each of five evaluated baselines on every task, increasing task-averaged return from 310.5 for the matched backbone to 715.5 for CaNI (approximately 130%). On the driving-video benchmark without jitter, time-averaged return improves by 15% with comparable final returns. Frozen-policy evaluations show improved robustness to evaluation-time jitter, and disjoint training and evaluation splits of the DAVIS video dataset show better performance on unseen backgrounds on two tasks. Posterior diagnostics and controlled ablations support reduced shift sensitivity and the intervention design.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.