acceptodds
Under review as a conference paper at ICLR 2027

FOCWM: Factorized Object-Centric World Model for Latent-Action Dynamics

Abstract

Existing object-centric world models typically decompose scenes into object slots, yet each slot remains a homogeneous representation in which geometry, shape, and appearance are entangled. Such entanglement requires inverse dynamics models to implicitly extract action-relevant cues from object representations, complicating latent-action inference. To address this challenge, we propose a Factorized Object-Centric World Model (FOCWM), which introduces a structured object representation together with adaptive latent-action inference for object-centric latent dynamics modeling. Specifically, we decompose each object slot into three complementary factors: spatial geometry, occupancy, and appearance. We further propose an action-oriented adaptive fusion mechanism that integrates these factors into object-wise representations for latent-action inference, allowing each factor to contribute adaptively according to the underlying dynamics. The inferred latent actions, together with the factorized object states, are then used to predict future object states. Experiments on synthetic and real-world datasets show that FOCWM consistently improves future prediction performance and downstream behavior learning compared with existing object-centric world models. The source code will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.