acceptodds
Under review as a conference paper at ICLR 2027

OneLA: Learning Many Worlds from Latent Actions

Abstract

Learning world models across diverse domains calls for a unified control space, yet action interfaces vary across domains and action annotations are often scarce. Latent actions learned directly from videos offer a promising way to establish such a space without explicit action supervision. However, learning shared latent actions across domains remains challenging, as variations in appearance, embodiment, and dynamics can entangle domain-specific visual cues with transferable transition structure. To this end, we introduce OneLA, a simple framework that learns continuous latent actions through a shared orthogonal dictionary and a common predictive objective, without task-specific designs or cross-domain alignment losses during latent action pretraining. These latent actions provide a shared conditioning interface for world models, with lightweight adapters mapping native control commands into the same space. Conditioned on latent actions and the current visual state, the world model predicts context-dependent transitions for controllable rollout. OneLA improves action readout across domains and reduces action cycle-transfer rotation error by 24.8% over the strongest evaluated baseline. A joint-training comparison further shows a 20.5% improvement in Human Preference Score v3 (HPSv3) for real-scene generation. These results suggest a path toward general-purpose world models that learn from diverse videos and support flexible control through latent action space.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.