CUDiT: Cubic Uniform Diffusion for Efficient Generation of High-Dimensional Discrete Visual Representations
Abstract
Discrete visual generation casts image synthesis as token prediction, providing a modeling interface shared with language and a natural path toward unified multimodal architectures. Recent visualtokenizers pursue larger effective vocabularies to reduce quantization error and higher-dimensional representations to retain the semantic information required for visual understanding. Conventional quantization methods face substantial challenges when scaling to high-dimensional representations and large codebooks. Dimension-wise quantization provides a simple and effective route to near-lossless quantization, but induces an enormous and highly redundant product space. The enormous code space makes direct joint categorical modeling intractable, so practical generation must rely on parallel factorization. The existing absorbing-diffusion approach uses large sampling budgets in this space and reports substantially lower quality at fewer steps. Our analysis shows that the resulting mismatch between the product of coordinate-wise marginals and the joint distribution becomes increasingly severe as fewer sampling steps require larger parallel updates. To address this problem, we propose CUDiT (Cubic Uniform Diffusion Transformer), a full-support discrete diffusion framework for efficient parallel generation of high-dimensional representation cubes. By denoising through continuous-time Markov chain dynamics over noisy token states, avoids the severe degradation that arises when a model trained only on clean tokens encounters the intrinsic noise of parallel generation. On ImageNet-256, -XL achieves gFID in steps, delivering comparable generation quality with up to two orders of magnitude fewer estimated backbone FLOPs than CubiD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.