acceptodds
Under review as a conference paper at ICLR 2027

Causality-Inspired World Action Model via Contrastive Representation Learning

Abstract

World-action models (WAMs) transfer predictive visual priors to robotic manipulation, but robust control requires invariance to task-irrelevant appearance changes that remain relevant to video prediction. Action supervision specifies the desired behavior for each observation without explicitly identifying which visually different observations share the same manipulation state. We introduce CausalWAM, a post-training method that supervises state correspondence through *action-preserving visual interventions*. We vary scene appearance while preserving robot–object geometry, task instructions, and demonstrated actions, changing the observation mechanism without changing the underlying manipulation state. A contrastive objective aligns alternative renderings of the same state and distinguishes different states within each task. A shared content reader connects this supervision to action learning, supplying compact tokens alongside the action expert’s native visual conditioning. The same design supports interface learning over a frozen video backbone and joint video–action adaptation. With action demonstrations restricted to Clean scenes, CausalWAM-Joint increases RoboTwin 2.0 Randomized success from 1.9% for FastWAM to 34.7%. It achieves 79.4% mean success across seven unseen LIBERO-Plus perturbation types and improves real-robot success under visual disturbance from 15.6% to 46.7%. Across 50 RoboTwin tasks, CausalWAM-Interface improves cross-appearance state retrieval by 9.42 percentage points. These results demonstrate the value of intervention-defined state correspondence for adapting predictive world representations to visually robust manipulation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.