acceptodds
Under review as a conference paper at ICLR 2027

OCCAM: Object-Centric Causal World Model via Vision-Language Graph Grounding

Abstract

Latent world models trained via joint-embedding predictive architecture (JEPA) excel at planning from pixels, but their transition functions are monolithic, lacking explicit representations of objects and interaction dynamics necessary for controllable and explainable embodied planning and reasoning. To address this issue, we introduce Object Centric CAusal World Model (OCCAM), a framework that integrates language-grounded causal structures directly into the transition function. OCCAM first uses open-vocabulary segmentation and tracking to obtain object-level latent states without hand-labeled masks. Then, OCCAM queries a vision-language model over motion-salient frames to build a sparse causal graph, validated by an independent conditional mutual information estimator. By combining language-grounded object abstraction with verified causal graph discovery, OCCAM enforces sparse, explainable state updates that isolate task-irrelevant entities, mitigating the extensive data requirements in classical causal discovery. We systematically evaluate OCCAM, analyzing dense, object-centric, and causal transition regimes across tasks of increasing visual and relational complexity. Experimental results demonstrate that OCCAM achieves a highly interpretable, object-centric latent space while preserving goal-conditioned planning accuracy. Crucially, on the complex multi-object LIBERO benchmark, OCCAM achieves superior downstream task control compared to dense and object-centric baselines, indicating that the structured causal reasoning is essential for reliable control as scene complexity scales.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.