acceptodds
Under review as a conference paper at ICLR 2027

Distilling Specialized Orders for Visual Generation

Abstract

Autoregressive (AR) image generators inherit the scalability and mature inference stack of large language models, yet most of them commit to a fixed raster-scan order. This choice precludes zero-shot inpainting, outpainting, and editing without retraining. Any-order AR models remove this restriction by learning to generate under arbitrary token permutations, but they must spread a finite capacity over factorizations and consequently trail fixed-order models in sample quality. We introduce Ordered Autoregressive (OAR) generation, a self-distillation procedure that resolves this tension. Starting from a pretrained any-order model, OAR extracts a specialized, content-dependent generation order for every training image from the model's own confidence scores, and then fine-tunes the model on these image–order pairs. Order specialization reallocates capacity from the full permutation space to a single high-confidence trajectory per image, while leaving the any-order conditionals, and hence the model's flexibility, intact. On class-conditional ImageNet , OAR improves FID from 2.39 to 2.17 over its any-order baseline under an identical evaluation protocol, and outperforms a random-to-raster annealing schedule applied to the same backbone (2.34). The same model performs zero-shot inpainting and outpainting and surpasses the bidirectional MaskGIT and MAR models on both tasks. It also retains multi-token parallel decoding. OAR generalizes to text-to-image generation with a different, scratch-trained backbone (FID on Fashion Products and on Multimodal CelebA-HQ), and human raters prefer its samples over the baseline in 64% of comparisons. On ImageNet, the procedure adds a fine-tuning stage of at most of the pretraining epochs, and it requires no architectural change or extra annotation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.