RoAR-VAE: Improving Latent Diffusability through Robust Prediction
Abstract
An effective visual tokenizer must preserve image fidelity while producing latents that downstream diffusion models can learn efficiently. Prior work improves latent diffusability through semantic alignment or joint tokenizer–diffusion training. Revisiting REPA-E, a representative approach, we find that removing its diffusion objective and dedicated generative blocks while retaining semantic prediction from corrupted latents yields comparable downstream generation performance under a matched setting. This observation motivates a general principle of *robust prediction*: generation-relevant information should remain recoverable from corrupted latents. We instantiate this principle with two complementary objectives. *RobustAlign* predicts pretrained semantic features from corrupted latents to preserve high-level structure, while *RobustRecon* reconstructs the clean image to retain fine-grained spatial and appearance information. Together, they form **Ro**bust **A**lignment and **R**econstruction VAE (**RoAR-VAE**). We attribute the resulting improvement in diffusability to robust prediction making useful information more accessible in high-variance latent directions, which diffusion models tend to learn earlier. Empirically, RoAR-VAE improves average generation benchmark performance by 2.19 points over the strongest baseline tokenizer under the same training budget, while improving small-text reconstruction accuracy by 16.93 percentage points over FLUX.1-VAE. These results demonstrate that RoAR-VAE achieves a favorable balance between reconstruction fidelity and downstream generative learnability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.