A Unified Rate–Distortion Perspective on Vector, Product, and Scalar Quantization
Abstract
Discrete visual tokenization relies mainly on vector, scalar, and product quantization, but lacks a unified conceptual framework for understanding their tradeoffs. In this paper, we propose a unified rate-distortion perspective on modern discrete visual tokenization. By viewing quantization as lossy compression, we characterize the nominal fixed-length coding rate through token count and codebook size, and quantization error as the distortion. Within this framework, we resolve three central questions. First, our theoretical and empirical results identify minimizing distortion, rather than maximizing codebook utilization, as the primary intrinsic objective for reconstruction fidelity and link distortion directly to the STE-induced gradient discrepancy. Second, we identify two conditions for fair intrinsic comparison: controlling latent feature statistics and matching nominal coding rates. Third, under these conditions, we recover the VQ-PQ-SQ distortion hierarchy in modern visual tokenization; modern VQ methods attain the lowest distortion in our experiments. The rate-distortion perspective clarifies quantizer evaluation by isolating intrinsic quantization effectiveness under controlled fixed-rate constraints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.