acceptodds
Under review as a conference paper at ICLR 2027

Tokenscale: Training-Free High-Resolution Image Generation In Token Space

Abstract

Diffusion Transformer (DiT)-based flow models achieve strong synthesis performance at training resolution, yet their extension to high-resolution (HR) generation frequently suffers from blurring and spatial layout collapse. We attribute this degradation to the token level. As resolution increases, image tokens grow quadratically while text tokens remain fixed, thereby diluting image–text attention and disrupting prompt alignment. In addition, persistent low-resolution reference guidance can restrict the synthesis of new HR details. To address these issues, we propose TokenScale, a training-free token-level framework. TokenScale interpolates complete image-token vectors obtained from low-resolution generation, rather than individual positions on the unpacked latent grid, and re-noises the expanded representation to initialize high-resolution sampling. During an initial denoising window, selective reference fusion corrects clean predictions in spatially coherent high-deviation regions, leaving unselected tokens unchanged and imposing no further reference correction afterward. This design supports structural consistency while allowing later denoising to synthesize details beyond the low-resolution content. Extensive experiments demonstrate that TokenScale substantially improves HR image generation quality, outperforms state-of-the-art training-free methods, and can be integrated with positional-encoding-based approaches for further gains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.