acceptodds
Under review as a conference paper at ICLR 2027

V-RAE: Rethinking Video Latent Spaces for Generation

Abstract

Latent video generation relies on autoencoders to define a compact space in which generative models operate. Although video autoencoders have evolved substantially, their latent spaces are still optimized primarily for pixel-level reconstruction and capture limited high-level semantics. A reconstruction-optimal latent space, however, need not be well suited to generative modeling. We propose V-RAE, a video representation autoencoder that builds temporally compact generative latents on top of frozen vision foundation models (VFMs). A lightweight temporal pooling module reduces temporal redundancy while preserving semantic structure, and a video decoder reconstructs videos from the compressed features. We evaluate V-RAE with four representative VFMs on video reconstruction, semantic probing, and class-conditional generation. V-RAE achieves 2.13 rFVD on K600, outperforming all evaluated large-scale pretrained video VAEs. Its latents retain substantially more semantic information than conventional video tokenizer latents. V-RAE achieves gFVD scores of 117.86 on UCF101 and 19.16 on K600, and matches Wan2.2 VAE's converged gFVD with up to 6× faster convergence. We further show that reconstruction quality alone is insufficient to characterize generative utility and introduce tFVD, a temporal-coherence diagnostic that correlates more strongly with downstream generation quality than rFVD. Beyond video generation, V-RAE also improves future video prediction on Cityscapes over the Wan 2.2 VAE latent space under matched prediction settings. Taken together, these results show that frozen semantic representations provide an effective basis for video reconstruction, generation, and predictive modeling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.