Distilling Drifting Transformers with Representation Autoencoders
Abstract
Despite the significant training acceleration and promising performance, effectively distilling models in Representation Autoencoder (RAE) latent spaces remains challenging. In this work, we argue that RAE is competent at high-quality one-step generation. We achieve 1.48 FID with only 16-epoch distillation on ImageNet 256 dataset, surpassing various state-of-the-art methods. To achieve this, we quantitatively study the geometrical behavior of different underlying latent spaces. We conclude that conventional trajectory-based distillation methods heavily rely on priors of plain teacher trajectories, while RAE incurs more complex trajectories with poor properties due to the highly anisotropic latent space. We introduce the recently proposed drifting field as the distillation methodology, which makes use of semantically rich RAE latents and provides supervision without relying on teacher trajectories. Bridging our Drift-RAE with previous generative paradigms, we propose several insightful modifications, including the first extrapolation-based guided sampling method for one-step generation with minimal additional overhead.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.