CanvasFlash: Efficient Autoregressive Image Generation with Semantic Representation
Abstract
Long visual sequences impose substantial training and decoding costs on high-resolution autoregressive (AR) image generation. We propose CanvasFlash, a novel approach that reduces this cost by exploiting the redundancy within visual representations. We depart from the reconstruction-oriented design in favor of a semantic image tokenizer that serves as a high-level information bottleneck, and delegate pixel rendering to a diffusion decoder. The resulting token distribution becomes highly concentrated: diverse tokens cluster in semantically informative regions, whereas broad regions share a common Canvas Token, indicating that only a fraction of the sequence is needed to characterize an image for semantic representation. Leveraging this sparsity, we propose Canvas-Sparse Training, which processes of the tokens required by full-sequence training. We also propose Canvas-as-Draft Decoding, which uses the Canvas Token as a parameter-free draft and achieves up to speedup in AR decoding. With reinforcement learning, CanvasFlash combines these efficiency gains with competitive image generation performance. More broadly, CanvasFlash highlights the redundancy carried by reconstruction-oriented visual representations, and validates semantic tokenization as a practical route toward efficient autoregressive image generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.