Noise-Aligned Autoencoders: Learning Diffusable Latents without End-to-End Training
Abstract
Diffusability measures how amenable an autoencoder's latent space is to efficient generative modeling, a property that often trades off against reconstruction fidelity. REPA-E improves the diffusability of autoencoders with an end-to-end training paradigm, but it is confined to controlled academic framework, i.e., class-conditional generation on explicitly labeled ImageNet. In this paper, we demonstrate that end-to-end tuning might not be essential, and the diffusability gain stems largely from representation alignment (REPA) rather than the diffusion loss. However, applying the REPA loss directly to the autoencoder can slip into a suboptimal solution that oversmooths the latent (e.g., VA-VAE), still lagging behind end-to-end tuned VAEs in the convergence speed of downstream diffusion training. Motivated by these observations, we propose Noise-Aligned Autoencoder (NA-VAE), which aligns the noise-conditioned latent with pretrained vision foundation models (VFMs) through a lightweight head with spatial inductive bias, forcing the latent to carry more low-frequency components that noise cannot drown out and thereby forming a structured latent space. NA-VAE is simple (implemented in < 4 lines of code) and efficient, requiring neither end-to-end tuning nor explicitly labeled data. Across VAE families, it consistently improves upon REPA and achieves performance comparable to or better than REPA-E on standard class-conditional and text-to-image benchmarks, while its decoupled training formulation offers potential for extension to video and audio generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.