ISI-WM: Interventional Selective Invariance for World Models
Abstract
Visual world models enable prediction and control directly from pixels, but reliable deployment requires representations that remain stable under irrelevant appearance changes without losing action-dependent dynamics. Existing robustness objectives typically align augmented views without identifying which variation is a nuisance or when distinct actions produce outcomes that must remain separable. We address this gap with Interventional Selective Invariance for World Models (ISI-WM), a training framework that suppresses background-induced variation while preserving distinctions associated with divergent action outcomes. A background intervention pairs observations of the same physical state rendered against different backgrounds and aligns their latent encodings. An action intervention forks alternative actions from a shared state, trains the latent rollout toward its executed branch, and activates separation and ranking only when the realized simulator outcomes diverge. Together, these signals optimize the existing RGB encoder and latent dynamics without extending the underlying architecture; simulator access is confined to training, and deployment remains RGB-only. Under a unified dynamic-background protocol, ISI-WM achieves stronger overall generalization to unseen backgrounds than the evaluated baselines, while ablations and latent-response analyses support the complementary roles of the background and action interventions. Code is available at https://anonymous.4open.science/r/isi-wm-code-E0AF.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.