Short Circuit VAE: Using Pre-Trained Variational Autoencoders For Latent Upscaling
Abstract
Scaling modern diffusion models to high resolutions introduces a computational bottleneck that grows quadratically () with the spatial dimensions and frequently degrades structural realism. Pixel-space super-resolution (SR) models inherently require heavy architectures to process uncompressed RGB data, and risk semantic drift when generating fine details from scratch. Existing latent upscalers either require slow, multi-step denoising loops, or rely on external architectures that ignore the pre-trained features of the base model, demanding extensive training and often producing smoothed outputs. To overcome these limitations, we introduce Short-Circuit VAE (SC-VAE). Instead of building a complex upscaler from scratch, SC-VAE acts as a shortcut directly inside the pre-trained Variational Autoencoder (VAE). Specifically, we reuse the early layers of the VAE decoder to extract features and spatially upscale the latents. These upscaled features are passed through a small, trainable adapter, and then fed into the final layers of the VAE encoder to project them back into the standard latent format. By relying on the VAE's own boundary layers to format the data, our method safely upscales images while remaining fully compatible with the generative model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.