acceptodds
Under review as a conference paper at ICLR 2027

RAE-VAR: Coarse-to-Fine Generation as Constrained Allocation in Representation Space

Abstract

Representation autoencoders provide semantic features for image generation, but how should these features be organized for coarse-to-fine prediction? Much of their decodable content can be retained in a low-rank channel subspace, yet placing that information in the earliest tokens can make those tokens difficult to predict. We introduce RAE-VAR, guided by constrained allocation: a design principle, rather than an explicitly optimized objective, for making early scales informative without asking them to predict too much from a short context. Its staircase tokenizer grows spatial resolution and channel rank together, using a nested PCA basis and Fourier resampling while leaving the VAR transformer unchanged. On ImageNet , the 310M-parameter, 16-layer model reaches gFID in 43 epochs, surpassing released VAR's after 200 epochs with fewer generator epochs. At 200 epochs, RAE-VAR improves to with nine next-scale sampling steps and 655 tokens. Larger 20- and 24-layer models approach their released VAR counterparts with roughly one fifth of the generator training epochs. Controlled ablations support both the structured basis and Fourier resampling, beyond what final reconstruction quality alone explains. Linear probes further show that early prefixes contain class information before their decoded images become recognizable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.