Unifying Latent Denoiser with Decoder Makes Robust Image Generation
Abstract
Despite the substantial progress of latent diffusion models, the gap between generated latents and those encoded from real data remains, posing a fundamental tension between reconstruction fidelity and generative modeling. In this study, we explore a simple-yet-effective approach that jointly learns denoising and image reconstruction from encoded image latents, termed the Denoise-decode Joint Transformer (DJT). During training, the denoising and reconstruction objectives shape the shared backbone together, while each task-specific head specializes in its corresponding task. At inference time, the model acts like a recursive transformer: it first operates as a denoiser to iteratively produce a final latent sample and then as a decoder to map the latent into an image. Experiments on ImageNet show that DJT-XL achieves gFID=**0.97**, with **30%** less parameters than the counterpart pipeline. Representation analysis further shows that the joint training strategy strengthens the alignment between denoising and decoding representations and improves their functional compatibility. The code and models will be publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.