DreamX: Video Diffusion Model is a Universal Latent Action Learner
Abstract
Learning a universal action representation from video requires modeling dynamics that remain consistent across different executions of an action while being invariant to changes in embodiment and appearance. Existing latent action models (LAMs) infer actions from visual transitions, but their representations often conflate the underlying action with irrelevant visual variations, which limits transfer across embodiments. We introduce DreamX, a framework that transforms a pretrained video diffusion model into a universal latent action learner. Our approach extracts compact latent actions from the model's rich spatiotemporal features, compressing each latent video frame by 300 while using its generative prior to disentangle action-relevant dynamics from visual appearance. The resulting latent space serves as a shared action interface across human and robot embodiments, environments, and action parameterizations. On LARY-Bench, DreamX achieves state-of-the-art representation quality. More importantly, actions encoded in one setting remain effective in another: the universal latent space supports any-to-any video transfer across embodiments and environments, as well as zero-shot cross-embodiment action transfer. This shared interface also extends beyond representation learning to control, where a compact latent policy improves task performance while lowering the computational cost of video-diffusion-based control.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.