Asymptotically Optimal Quantizer-Guided Sinkhorn Alignment for Mitigating Codebook Collapse
Abstract
Codebook collapse remains a fundamental limitation of vector quantization (VQ), causing poor codebook utilization and increased quantization error. While recent methods improve utilization through distribution alignment, directly matching the codebook to the encoder distribution can deviate from the asymptotically optimal codeword distribution, leading to excessive codeword concentration in high-density regions and insufficient coverage of lower-density regions. We introduce AOptSA, a theory-driven quantization framework motivated by asymptotically optimal quantizer theory. Its key insight is to align the codebook with a density-reweighted target distribution, rather than the raw encoder distribution, thereby better balancing probability mass and spatial coverage. AOptSA realizes this target online via importance sampling from encoder outputs and minimizes Sinkhorn divergence between the sampled target and codebook distributions. The resulting alignment adapts throughout training without altering nearest-neighbor quantization or latent dimensionality. Across image benchmarks, AOptSA achieves full codebook utilization, consistently improves reconstruction quality over strong baselines, and maintains its gains as codebook size increases.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.