acceptodds
Under review as a conference paper at ICLR 2027

State Before Pixels: Large-scale Text-to-Video Generation in Representation Space

Abstract

Video representations learned by vision foundation models encode scene semantics and dynamics as spatiotemporal world states beyond pixel appearance. While representation autoencoders have shown promise for image generation, representation-space video generation remains largely underexplored, with their higher dimensionality posing additional challenges. In this paper, we demonstrate that native video representations can serve as effective modeling targets for large-scale, open-domain text-to-video generation. Our approach, **WorldFlow**, separates world-state modeling from visual rendering: a text-conditioned flow model first generates latent world states directly in native V-JEPA space, and a second-stage flow model then renders these states into high-quality videos in VAE space. This design retains the learned semantic structure of pretrained video representations as the basis for generation, while allowing fine visual details to be synthesized in a space designed for reconstruction. To enable effective diffusion-transformer denoising in this high-dimensional representation space, we systematically study its geometry and prediction parameterization. We find that by conducting Riemannian Flow Matching on the native representation manifold with *v*-prediction, high-dimensional video representations can be effectively modeled without introducing additional compression. Extensive experiments demonstrate that WorldFlow achieves high-quality text-to-video generation. Beyond RGB, the same generated world states can be readily rendered into well-aligned multimodal outputs such as depth, segmentation, human pose, *etc*. These results demonstrate the feasibility of generating unified world states containing appearance, semantics, geometry, and dynamics within a shared representation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.