Aligned Geometry Latent Space for Improved World Consistency in Joint Video-Geometry Modeling
Abstract
While jointly modeling video with compressed geometry latents has been shown to effectively improve geometric consistency in video generation, such methods typically suffer from slow convergence or training instability. Our analysis suggests that this issue is closely related to the misalignment between compressed geometry latents and pretrained video models: higher reconstruction fidelity does not necessarily make the latents more effective for video models to understand and utilize. Motivated by this, we propose geometry latent alignment (GeoLA), a post-training framework that directly leverages the video generator to supervise geometry compression and produce a well-aligned geometry latent. Specifically, we use geometry latents compressed by a geometry adapter as conditions for video denoising, and optimize the adapter with the diffusion loss. This allows the video model to implicitly guide the adapter to retain the geometric information most useful for denoising. We then freeze the aligned adapter for joint video-geometry modeling, followed by dual-stream distribution matching distillation for low-latency joint generation. To support training and evaluation, we further introduce GeoVid-80K, a dataset containing over 80,000 video clips with rich geometric variations. Experiments show that GeoLA converges faster during joint training and generates videos with improved geometric consistency.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.