acceptodds
Under review as a conference paper at ICLR 2027

Learning Spatial-Pyramid Visual Representations for Autoregressive Image Generation

Abstract

Autoregressive image generation relies on two coupled components: visual token representation and token distribution modeling. Existing tokenizers either preserve spatial correspondence through single-scale grids or aggregate global information with flat sequences, without explicitly organizing tokens across spatial granularities. Although end-to-end training aligns representation learning with generation objectives, it does not by itself determine how hierarchical visual information should be organized for conditional prediction. We introduce a spatial-pyramid 1D visual representation that organizes global tokens and regional tokens at multiple spatial granularities within a single 1D sequence. The corresponding context model generates global tokens causally and predicts each regional level in parallel, conditioned on all coarser groups, yielding a global-to-local factorization. This factorization retains all latent tokens while reducing sequential decoding depth compared with token-wise prediction. We jointly optimize the tokenizer and context model so that the learned representation adapts to the proposed factorization. On ImageNet at 256256 resolution, our method achieves a gFID of 1.21 with classifier-free guidance using 19 sequential decoding steps.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.