Image Tokenizers Should Be Trained Between Images
Abstract
A visual tokenizer encodes an image into a smaller representation, and its decoder turns that code back into the image. However, the decoder is trained only at the codes the encoder produces from the training images. It never learns what to do with codes that lie in the space between training samples. Codes in this space lie outside of the training distribution and the decoder produces blurry images with strong artifacts there. For image generation, a second model, the generator, learns to produce images as codes of the encoder of a tokenizer. But, if generated codes land outside of what the decoder expects, the decoded image will contain imperfections. In order to make the decoder more robust to generated codes, current methods add noise to codes during decoder training. However, recent research indicates that generated codes do not simply contain noise. Instead, they lie between the codes the generator was trained on, and resemble an average of codes of training images more than one of them plus noise. Therefore, to improve quality of generated images, decoders should be trained to handle codes in the space between training samples. Motivated by these insights, we propose InterpTok, which fine-tunes MAETok by training the decoder to generate natural-looking images on mixtures of the codes the encoder produces. But no target image exists for an image between two training samples. So a discriminator judges whether each decoded mixture of codes looks as natural as the images whose codes were mixed. Additionally, a mixing-ratio loss requires the decoded image to visually keep the mixing ratio of the two codes. This prevents a collapse to a trivial solution and trains the decoder to produce a sharp and natural composition of both images. InterpTok keeps MAETok's reconstruction quality and improves interpolation quality more than threefold, the best among the thirteen tokenizers of a released comparison. Beyond interpolation, we show that our method can substantially improve generation training speed as well as the generation quality with much less computational cost. InterpTok improves the generation FID of a matched diffusion transformer from 11.61 to 8.48 without guidance, and even reaches 8.11 at an eighth of the sampling steps, where the generator on MAETok's codes worsens. Furthermore, the generator trained on our tokenizer surpasses the final results of the generator trained on MAETok at half the training steps. Our findings show that the space between image codes can and should be explicitly supervised rather than ignored.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.