VideoRAE: Taming Video Foundation Models for Generative Modeling via Representation Autoencoders
Abstract
Video generation models typically rely on 3D-VAEs trained for pixel-level reconstruction, whose latent spaces may underrepresent semantic structure. We introduce VideoRAE, a representation autoencoder that converts features from a frozen video foundation model into compact, reconstruction-capable latents for video generation. Lightweight projectors compress multi-scale hierarchical features into compact 1D sequences or structured 3D grids, supporting continuous latents for Diffusion Transformers and discrete tokens for autoregressive models through multi-codebook high-dimensional quantization. During decoding, a local–global representation alignment objective transfers semantic structure from the frozen encoder and removes the need for KL regularization. Comprehensive experiments show that VideoRAE achieves strong reconstruction in both continuous and discrete regimes. On UCF-101, autoregressive and diffusion generators built on VideoRAE achieve class-conditional gFVD scores of 40 and 93, respectively, while converging approximately 5x faster than autoencoder baselines. In a 4B-scale text-to-video study with randomly initialized diffusion backbones, frozen VideoRAE-3D achieves higher VBench total scores than Wan2.1-VAE at all 11 evaluated checkpoints, reaching 69.07 versus 68.24 at 200k training steps. These results establish frozen video foundation representations as compact, versatile, and generation-friendly video latents. Code and models will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.