acceptodds
Under review as a conference paper at ICLR 2027

Frequency-Steered VAEs: Restoring Fine Details in Low-Resolution Latents

Abstract

Latent diffusion models (LDM) rely on VAEs to compress visual inputs for computational efficiency, which is particularly consequential for very high resolution visual inputs. Through an in-depth interpretability analysis, we show that this compression disproportionately degrades high-frequency information at lower spatial resolutions, discarding fine details such as facial features or text patterns. Building on this insight, we propose a we call that uses a higher-resolution VAE as a teacher to supervise its lower-resolution counterpart, recovering lost frequency information via a band-wise . FS-VAE enables richer reconstruction from low-resolution latents without expanding channel dimensions, reference encoders, external models, or retraining downstream diffusion architectures. Furthermore, we analyze the impact of distinct loss objectives on reconstruction quality and demonstrate our approach's generalizability across diverse image and video VAEs. By preserving the original architecture and latent dimensionality, FS-VAE acts as an efficient, plug-and-play replacement for existing LDM pipelines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.