acceptodds
Under review as a conference paper at ICLR 2027

MoFa-World: Motion-Factorized Latent Actions for Interactive World Models

Abstract

Egocentric world models provide a promising foundation for interactive environment simulation, yet existing approaches tend to focus on egomotion control, with object-interaction control constrained by the scarcity of densely annotated interaction videos. Although latent action models (LAMs) learn action representations from videos without action labels, they often entangle object interaction and egomotion in a single representation. We believe that pixel-space motion patterns can reveal distinct action semantics, providing a basis for learning decoupled representations. Building on this insight, we introduce MoFA-LAM, which pretrains region-level latent actions through pixel-space motion factorization and aggregates them into interaction and egomotion representations through region-specific reconstruction with limited mask supervision. Probing experiments demonstrate the functional specialization of the two representations and the benefits of motion-factorized pretraining over vanilla LAM pretraining. We further introduce MoFA-WORLD, an egocentric world model conditioned on these representations through distinct pathways. It supports cross-scene action transfer and adaptation to dataset-specific controls, including joint use of interaction and egomotion controls learned from different datasets. Together, these results support pixel-space motion factorization as a way to learn decoupled control representations for egocentric world models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.