acceptodds
Under review as a conference paper at ICLR 2027

Learning Discrete Autoregressive Priors with Wasserstein Gradient Flow

Abstract

Discrete image generation commonly uses two-stage training: a tokenizer is first trained for reconstruction, and a prior is then fitted to the tokens from the frozen tokenizer. This decoupling leaves the tokenizer unaware of the model that will later generate its tokens. As a result, the learned tokens may preserve image information well but still be difficult for an autoregressive (AR) prior to predict from left to right. We analyze this mismatch using Tripartite Variational Consistency (TVC), which decomposes latent-variable learning into three consistency conditions: conditional-likelihood consistency, prior consistency, and posterior consistency. TVC shows that two-stage training preserves the reconstruction side but leaves prior consistency outside the tokenizer objective: the overall token distribution is fixed before the AR prior participates in training. Motivated by this view, we add a distribution-level prior-matching signal during tokenizer training, while keeping the reconstruction objective unchanged. We optimize this signal with a Wasserstein-gradient-flow update. For hard categorical tokens, we approximate the update with a token-level contrast between an auxiliary AR model that tracks the tokenizer's current token distribution and a frozen AR reference. Computing this tokenizer update requires only forward passes through the two AR models; the proxy uses its own cross-entropy loss to track the current tokens. After tokenizer training, a separate AR prior is trained for generation. The resulting tokenizer, wAR-Tok, reduces AR loss and improves generation FID on CIFAR-10 and ImageNet at comparable reconstruction quality.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.