acceptodds
Under review as a conference paper at ICLR 2027

DC-Gen 2.0: Efficient and High-Fidelity Image Generation in Deeply Compressed Latent Spaces

Abstract

Existing diffusion models excel at generating high-quality images, but scaling them to higher resolutions remains computationally expensive. Deep-compression autoencoders reduce latent token counts, yet maintaining reconstruction fidelity and synthesizing fine details in the compressed latent space remain challenging. To address these challenges, we present DC-Gen 2.0, built on our DC-AE 2.0 family. For DC-AE 2.0-Structured, we propose Pyramid-Aligned Decoder Training to facilitate downstream generation while maintaining high reconstruction quality. Building on the structured latent space introduced by DC-AE 1.5, it aligns channel budgets with hierarchical decoder paths and Gaussian-pyramid targets, encouraging coarse-to-fine reconstruction across decoder stages. To improve generation quality under high compression, we introduce Deep-Compression Detail Enhancement, which combines targeted data curation with generation post-training through diffusion adaptation and reinforcement learning. Curated text and face data support fine-detail preservation during autoencoder training and synthesis during diffusion post-training. Experiments show that DC-Gen 2.0 improves reconstruction quality and enables efficient, high-quality image generation under high compression. For FLUX.2-dev, DC-Gen 2.0 f32 achieves 3.39x speedup at 1024 and DC-Gen 2.0 f64 achieves 53.91x speedup at 4096.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.