acceptodds
Under review as a conference paper at ICLR 2027

Can Discrete Tokens Serve as Semantically Lossless Representations of Visual Signals?

Abstract

Discrete visual tokens offer a native interface for multimodal modeling under a shared next-token prediction objective by representing vision in a discrete form analogous to language. However, quantization error is widely assumed to cause unavoidable semantic loss, fundamentally limiting the semantic fidelity of discrete visual tokens and thereby positioning continuous visual representations as an intrinsic upper bound. We revisit this assumption using multimodal understanding as a proxy for semantic fidelity and find that quantization error is unreliable, as discrete representations with substantially larger errors can achieve better understanding. Instead, the effective bit budget of discrete indices, reflecting the semantic information they convey, emerges as a key determinant of downstream performance. However, existing quantizers eventually saturate as additional codebooks contribute progressively less semantic information. As one approach to alleviating this saturation, we introduce Factorized Vector Quantization, which quantizes complementary, fixed-dimensional views of the same visual feature with symmetric parallel codebooks. It continues to improve at larger bit budgets, ultimately matching continuous visual representations in multimodal understanding. We train on industrial-scale multimodal data and achieve state-of-the-art performance among discrete-representation models, improving over UniAR and Emu3-Chat on a nine-metric average by 22.8% and 59.2%, respectively. To the best of our knowledge, our model is also the first discrete-representation model to perform competitively with strong continuous-representation models such as Qwen3-VL, including on fine-grained OCR tasks. Beyond these results, we unveil the key principles and a practical recipe that make such competitiveness possible. These results and practical insights jointly support the potential of discrete visual representations to underpin large-scale native multimodal foundation models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.