acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Representation Alignment in Autoregressive Image Generation

Abstract

Representation alignment has improved image generation in both diffusion and autoregressive (AR) models by supervising intermediate latents with features from pretrained visual encoders. Nevertheless, in AR models, the utility of representation alignment depends on the causal position of the supervised state. Later states observe more image content but can influence fewer subsequent predictions, so higher representation fidelity does not necessarily improve generation. We introduce Causally Accessible Representation Alignment (CARA), grounded in a theoretical and empirical investigation of how visual representation learning can support future predictions in autoregressive generation. Our analysis indicates that applying alignment before image generation encourages learning visual information predictable from conditions. Experiments comparing supervision positions and depths support aligning the final-layer representation of a learnable prefix token with a global visual teacher target. The token's representations remain accessible throughout decoding, providing semantic context for subsequent predictions within the original generation pipeline. Experiments demonstrate improved class-to-image generation and text-to-image prompt adherence across the evaluated model sizes. Controlled comparisons support combining deep alignment with a dedicated prefix state, while cache interventions establish the state's predictive utility beyond the first image token. These findings position causal role as a design principle for representation alignment and motivate learning representations according to both the information they can encode and the future predictions they can support.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.