Frequency-Steered VAEs: Restoring Fine Details in Low-Resolution Latents
Abstract
Latent diffusion models (LDM) rely on VAEs to compress visual inputs for computational efficiency, which is particularly consequential for very high resolution visual inputs. Through an in-depth interpretability analysis, we show that this compression disproportionately degrades high-frequency information at lower spatial resolutions, discarding fine details such as facial features or text patterns. Building on this insight, we propose a we call that uses a higher-resolution VAE as a teacher to supervise its lower-resolution counterpart, recovering lost frequency information via a band-wise . FS-VAE enables richer reconstruction from low-resolution latents without expanding channel dimensions, reference encoders, external models, or retraining downstream diffusion architectures. Furthermore, we analyze the impact of distinct loss objectives on reconstruction quality and demonstrate our approach's generalizability across diverse image and video VAEs. By preserving the original architecture and latent dimensionality, FS-VAE acts as an efficient, plug-and-play replacement for existing LDM pipelines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.