Tokenscale: Training-Free High-Resolution Image Generation In Token Space
Abstract
Diffusion Transformer (DiT)-based flow models achieve strong synthesis performance at training resolution, yet their extension to high-resolution (HR) generation frequently suffers from blurring and spatial layout collapse. We attribute this degradation to the token level. As resolution increases, image tokens grow quadratically while text tokens remain fixed, thereby diluting image–text attention and disrupting prompt alignment. In addition, persistent low-resolution reference guidance can restrict the synthesis of new HR details. To address these issues, we propose TokenScale, a training-free token-level framework. TokenScale interpolates complete image-token vectors obtained from low-resolution generation, rather than individual positions on the unpacked latent grid, and re-noises the expanded representation to initialize high-resolution sampling. During an initial denoising window, selective reference fusion corrects clean predictions in spatially coherent high-deviation regions, leaving unselected tokens unchanged and imposing no further reference correction afterward. This design supports structural consistency while allowing later denoising to synthesize details beyond the low-resolution content. Extensive experiments demonstrate that TokenScale substantially improves HR image generation quality, outperforms state-of-the-art training-free methods, and can be integrated with positional-encoding-based approaches for further gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.