acceptodds
Under review as a conference paper at ICLR 2027

Causally Debiased Latent Action Model for Embodied Action-Conditioned World Models

Abstract

Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from videos without executable action labels, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual confounders. With such objectives, LAM favors visual context over action dynamics, therefore downstream ACWMs exhibit residual motion under zero actions and fail to reproduce supplied dynamics. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three debiasing objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and well-calibrated representations. CD-LAM reduces action-following error by up to 42% and 35% when conditioned on latent and robot actions, while improving fidelity, and matching DreamDojo reference with over 12 fewer adaptation updates. CD-LAM also adapts LTX-2.3-22B into an ACWM with competitive quality using only 3.2% of DreamDojo-14B's sample exposures. On X-VLA, pretraining with CD-LAM instead consistently improves success rates by up to 5% across LIBERO, LIBERO-Plus, and RoboTwin C2R.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.