How Much of Your Codebook Survives Denoising?
Abstract
Two-stage image generation compresses an image into discrete codewords—a tokenizer with codebook size K—then models the code sequences. How large should K be? Reported optima span two orders of magnitude. We treat the trained second-stage generator as a measurable noisy channel: renoising and re-denoising real latents yields its per-token confusion kernel, and a surrogate computed from the codebook alone reproduces the measurement within its validity domain. Three rates must be distinguished—nominal (log₂K), used (H(q)), and transferred through the probe—and three findings follow. Snapping: diffusion outputs land exactly on the codeword grid, so generation error means committing to the wrong codeword, invisible to dispersion-based diagnostics. Saturation: across CIFAR and ImageNet (64–256px) the nominal rate keeps growing while probe transfer stalls; in the converged CIFAR-10 sweep, about 8 of 15 nominal bits survive our stress probe. The saturation point, computed from the codebook alone, bounds the useful range of K before any stage-2 is trained: the best gFID inside this range matches the global best on both CIFAR variants and trails it by 0.1 on ImageNet. Allocation: in our matched-budget experiments, the catastrophic large-K collapse of autoregressive models is driven by vocabulary crowding out core parameters rather than by K itself, while diffusion failures trace to stage-2 commitment capacity or codebook geometry—interventions on each remove them. All numerical predictions were archived before measurement; five falsified predictions are reported.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.