acceptodds
Under review as a conference paper at ICLR 2027

Raster or Random? Let the Data Decide

Abstract

Autoregressive (AR) and masked diffusion models differ fundamentally in generation order, yet what determines the best order remains unclear. We show that the answer depends on the inner structure of the data. Prior comparisons contrast complete systems that differ in attention pattern, training objective, and decoding schedule at once, so the effect of order alone cannot be isolated. We introduce a unified framework that makes order the only variable: a shared causal-attention architecture takes the token order as an explicit input, holds everything else fixed, and decouples the training order from the inference order. Linear dynamical systems (LDS) and their shuffled counterpart further provide controlled proxies whose dependencies are aligned or misaligned with the token order. Across modalities, the AR order wins on language and ordinary LDS, whereas the random order wins on images and shuffled LDS. To reveal the underlying mechanisms, we decouple training and inference stages and reveal two distinct mechanisms: during training, order diversity acts as data augmentation on image-like data but not on language; during inference, scattered orders prevent the accumulation of shift along the generation. An information-theoretic analysis on LDS explains how the augmentation benefit depends on the data dependency structure. These results clarify when and why each order succeeds and guide the choice of order for a given modality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.