Just Image Transformers with Late-Stage Pixel Refinement
Abstract
Recent work on latent-free diffusion transformers has demonstrated the feasibility of training image generators without a latent VAE. This avoids compression artifacts from the VAE and simplifies the training pipeline. In particular, JiT showed that even a standard transformer can be trained to generate images when the model is trained to predict the final image (-prediction). Even though JiT unlocks training efficiency comparable to latent models, it still suffers from reduced image quality likely arising from modeling the image as a collection of large patches. In this work, we revisit JiT and show that adding simple convolutional pixel-pathway layers in a post-training phase significantly improves generation quality. Moreover, we show that the added pixel-pathway does not need to be active in the early denoising stages, and can be activated as late as in the last 20% of the denoising process. Our experiments on ImageNet@256 show improvements compared to the original JiT results, achieving an FID of , without using additional losses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.