DynamicTok: Dynamic Tokenization for Flexible Image Representation and Generation
Abstract
Existing image tokenizers typically use a fixed number of tokens at fixed spatial locations, overlooking the substantial variation in visual complexity both across images and within individual images. To address this limitation, we introduce DynamicTok (Dynamic Image Tokenizer), which adaptively allocates tokens in both quantity and spatial location according to content complexity. DynamicTok first predicts a heatmap that estimates regional importance, assigning more tokens to visually informative regions and fewer tokens to redundant ones. To process these spatially non-uniform tokens effectively, we propose a hybrid 1D–2D tokenizer in which latent tokens capture global context while receiving localized cues from image tokens. DynamicTok can be readily integrated with existing generative models. Moreover, we introduce a two-stage generation strategy that predicts a heatmap before image synthesis, enabling flexible generation under different token budgets. On ImageNet-256 reconstruction, DynamicTok achieves rFID scores of 0.57 and 0.39 with 64 and 128 tokens, respectively, outperforming prior methods under the same token budgets. For ImageNet-256 generation, DynamicTok attains a gFID of 1.40 with 64 tokens, establishing a new state of the art among methods using fewer than 256 tokens. Moreover, both our tokenizer and generator can dynamically allocate tokens in terms of number and spatial locations as needed.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.