Improving Reconstruction of Representation Autoencoder
Abstract
Recent work leverages vision foundation models as image encoders for latent diffusion models (LDMs), as their semantic representations are favorable for generative modeling. However, these representations often lack low-level visual information (\eg, color and texture), limiting reconstruction fidelity and posing a bottleneck to further scaling LDMs. To address this limitation, we propose LV-RAE, a representation autoencoder that augments pretrained semantic features with complementary low-level information, enabling high-fidelity reconstruction while preserving semantic structure. Beyond reconstruction, we identify a decoder-side source of the reconstruction–generation trade-off: sensitivity to fine-grained latent variations supports faithful reconstruction but can also amplify errors in generated latents and degrade generation quality. Our analysis suggests that excessive decoder responses along off-manifold directions contribute to this issue in high-dimensional latent spaces. We improve decoder robustness through noise-augmented fine-tuning and further enhance generation quality through controlled noise injection at inference. Moreover, scaling only the decoder, while keeping the latent representation and diffusion model fixed, simultaneously improves reconstruction and generation, increasing PSNR from 30.91 to 32.09 dB and reducing unguided gFID from 2.42 to 1.86. These results demonstrate that LV-RAE supports both high-fidelity reconstruction and strong generation, while highlighting decoder robustness and capacity as key factors governing the reconstruction–generation trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.