GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation
Abstract
Visual generators can produce photorealistic frames without preserving the scene they depict. We argue that this is not only a modeling problem but a representation problem: generators evolve latents organized around appearance, while perception systems recover geometry in a different space. Rather than add geometry as another output, we make a compact reparameterization of a geometry foundation model's feature space the generator's native state. We realize this shift with the geometry-native autoencoder (GAE), whose latent is jointly decodable to appearance, depth, cameras, and point maps. With this state, a standard conditional flow supports diverse generation tasks. In controlled comparisons that hold the generator and training protocol fixed, replacing the latent with GAE improves both visual quality and independently measured 3D coherence: FVD falls by and on RealEstate10K and DL3DV, and camera-trajectory error is halved on RealEstate10K. Together, these results show that the latent space is central to geometry-consistent generation and can serve as a shared interface between perception and generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.