Learned Image Compression with Vector Quantization-Based Prior Modeling
Abstract
Entropy modeling is essential for learned image compression (LIC), where the latents are usually scalar-quantized and subsequently encoded based on estimated priors. Compared to scalar quantization (SQ), vector quantization (VQ) is theoretically more efficient for compression performance. However, the codebook size in VQ grows exponentially with the bitrate, resulting in an extremely high-dimensional optimization landscape. Such a complex landscape poses significant challenges to convergence and often leads to severe codebook collapse, making it difficult to achieve effective rate-distortion (RD) optimization. To make VQ tractable within an end-to-end framework, we propose Vector Quantization-Based Prior Modeling (VQPM). VQPM applies VQ to a subset of the latents with the lowest spatial resolution. These low-resolution features inherently capture the image's global structure and are coded first, without any decoded context. VQ is particularly suitable here as it can model the dependencies within the vector itself, compensating for the lack of external context. The vector-quantized latents then serve as global structural guidance for the subsequent hyperprior and contextual prior modeling. We further introduce an RD-oriented codebook training strategy that optimizes the codebook using end-to-end RD gradients. Experimental results show that VQPM achieves state-of-the-art RD performance, reducing PSNR BD-rate over VTM-22.0 by 20.29%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.