Dream-Lift-VLA: Generative 2D-to-3D World Dreaming for Vision-Language-Action Models
Abstract
Vision-language-action (VLA) models offer a promising path toward generalist robot policies. A key challenge is that sparse action supervision provides limited guidance for capturing fine-grained future scene changes. To address this limitation, recent methods incorporate world modeling into VLAs by jointly predicting future 2D videos and 3D geometry. However, these two forms of supervision are inherently ordered, with future 2D prediction first capturing current-to-future scene evolution within the same modality and future 3D prediction subsequently grounding this evolution across modalities in 3D space. Parallel training ignores this order by imposing 3D grounding before the 2D scene evolution has been sufficiently learned. To address this issue, we propose **Dream-Lift-VLA**, a progressive VLA learning framework that includes two phases: **Dream** and **Lift**. Specifically, the Dream phase uses world embeddings to condition a pretrained video diffusion model for future-video generation, enabling the VLA to learn fine-grained scene dynamics and manipulation semantics. Building upon these learned dynamics, the Lift phase uses the world embeddings to condition a pretrained depth diffusion model for future-depth generation, further shaping the VLA representations with 3D spatial awareness. Empowered by this progressive 2D-to-3D representation learning, our policy develops stable representations of both manipulation semantics and 3D geometry. Extensive experiments on diverse simulation and real-world tasks demonstrate that Dream-Lift-VLA consistently outperforms state-of-the-art approaches. The code of this work will be released soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.