Token Allocation For Better Generation
Abstract
Autoregressive image generation relies on visual tokenizers to convert images into discrete tokens. Recent variable-rate tokenizers dynamically allocate different token lengths to different images to optimize reconstruction. In this paper, we first show that token allocation optimized for a reconstruction objective, such as mean-square-error (MSE), might lead to worse generation. Next, we propose (Generation-Aware Token Allocation), a variable-rate tokenization method optimized for generation. More specifically, we first introduce a latent-context formulation for variable-rate autoregressive modeling to enable generation with flexible token lengths. Further, we propose a selection criterion that assigns different token lengths to images for better generation. Empirically, GTA consistently improves downstream generation across different tokenizers. On GigaTok, GTA reduces gFID from 5.50 to 4.75 compared to its fixed-length counterpart.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.