LatentWeave: Are Semantics Missing from VAE Latents—or Merely Implicit?
Abstract
Visual encoders define the representation spaces underlying visual reconstruction, understanding, and generation. Reconstruction-oriented VAEs provide compact representations with strong visual fidelity, yet their latents are shaped primarily by reconstruction and exhibit weak semantic organization. Recent approaches address this limitation by replacing VAE encoders with pretrained representation encoders. We revisit whether such replacement is necessary and find that substantial semantic information remains recoverable from pretrained VAE latents despite being poorly exposed in their native organization. Motivated by this observation, we introduce LatentWeave, a unified framework for constructing and generatively modeling semantically organized visual latents from pretrained VAE representations. It comprises WeaveAE, which organizes complementary detail and semantic strands, and WeaveDiT, which explicitly exploits this dual-strand structure during generative modeling. We further introduce Self-REPA, which reuses WeaveAE’s semantic representation as an internal alignment target to improve generative learning without introducing an external vision encoder during generator training. Across image and video benchmarks, LatentWeave improves semantic accessibility and generative learning efficiency while preserving reconstruction fidelity. On UCF-101, WeaveAE achieves an rFVD of 3.11 for reconstruction, while WeaveDiT achieves an FVD of 135.49 for class-conditional generation. Our results establish latent reorganization from pretrained VAE representations as a practical paradigm for unified visual representation learning and efficient generative modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.