Attentive Vector Quantization
Abstract
Discrete latent representations are widely used in generative and representation learning, but conventional vector quantization commonly relies on hard nearest-neighbor assignment with a straight-through estimator (STE), introducing a mismatch between forward computation and backward gradients. Learned vector quantizers can also exhibit highly non-uniform code usage or codebook collapse. We introduce Attentive Vector Quantization (AVQ), an STE-free differentiable quantization approach that replaces nearest-neighbor assignment during training with temperature-controlled cosine-similarity attention between latent tokens and a learned codebook. AVQ combines a soft-to-hard training curriculum, a compact codebook-usage regularizer that encourages confident local assignments and balanced global utilization, and hard-code decoder calibration to bridge differentiable training and discrete inference. We evaluate AVQ against conventional vector quantization (VQ), EMA-based vector quantization (VQ-EMA), and finite scalar quantization (FSQ) on COCO, Tiny ImageNet, and MiniPlaces. Across all three datasets, AVQ achieves complete codebook utilization with near-maximal hard-code perplexity and the lowest rFID among the evaluated single-stage quantizers. On COCO, across three model seeds, AVQ uses all 1024 codes with a perplexity of and achieves an rFID of , an 18.0% reduction relative to the best baseline mean. Similar improvements in rFID are observed on MiniPlaces and Tiny ImageNet, while AVQ maintains complete utilization and perplexity above 1000 on all three datasets. Across all single-stage AVQ ablations, complete codebook utilization persists under changes in codebook size, embedding dimension, normalization, learned code scaling, and spatial rate. We further extend AVQ to Attentive Residual Vector Quantization (ARVQ). Across all three datasets, ARVQ maintains near-maximal stage-wise perplexity above 1016 at every residual stage, whereas EMA-based residual vector quantization (RVQ-EMA) exhibits substantially lower and less consistent stage-wise perplexity. These results indicate that attentive quantization provides robust use of discrete capacity in both single-stage and residual quantization settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.