Hybrid-Diffusion: Latent-to-Pixel Diffusion for Limited Data Generation
Abstract
Pixel diffusion models have achieved remarkable success in synthesizing fine textures and local structures by denoising score matching. Despite this success, learning accurate score functions in such a high-dimensional space requires substantial data, making it difficult to reliably capture global semantic and structural information under limited data regimes. Latent diffusion addresses this challenge by removing pixel-level redundancy and performing diffusion in a compressed representation that retains salient semantic and structural information. However, because the autoencoder is typically frozen and lossy, this compression can limit the fidelity of fine-grained details. In this paper, we introduce Hybrid Diffusion, a hybrid latent-to-pixel diffusion framework that effectively combines the advantages of both pixel and latent diffusion models for limited data generation. Specifically, Hybrid Diffusion consists of both latent and pixel stages, and performs most of the early denoising in the latent stage before switching to the pixel stage at a chosen transition point, thereby substantially narrowing the pixel-space learning problem while addressing the representation bottleneck of latent space. Importantly, we make the pixel stage explicitly aware of latent prediction errors by using decoded clean predictions from randomly forward-corrupted training latents as training conditions, allowing the pixel stage to address these errors alongside autoencoder reconstruction artifacts. Experiments across multiple datasets demonstrate consistent improvements over both pixel and latent diffusion baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.