Exploring the Design Space of VAEs for Diffusion-Based Image Generation
Abstract
Latent diffusion models often retain VAE encoder architectures introduced by early LDM systems, despite significant evolution in both decoders and diffu- sion backbones. We revisit the encoder from an efficiency perspective, and study how the computation should be allocated with regard to operations and resolution stages. We introduce ABCRP and LatentMaid: encoders built from pixel unshuffle, RMSNorm, and a depthwise + SwiGLU residual block. The design keeps spatial effects in the decoder and uses only local, lightweight components in the encoder, explicitly restricting cross-spatial dependencies in the encoder without requiring an auxiliary locality loss. At CFG=1, ABCRP with standard residual blocks achieves the best generation quality in our study (17.68 gFID), while LatentMaid achieves 18.32 gFID with 89% fewer encoder FLOPs than the SD-VAE baseline and increasing measured encoder throughput by approximately 3×, providing a substantially more efficient operating point. All code has been uploaded as supplementary material - available for testing and reproduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.