MBVQ: Multi-Level Block-Wise Vector Quantization for Hardware Friendly Post-Training Quantization of LLMs
Abstract
The rapid scaling of large language models (LLMs) has substantially increased their deployment cost. Post-training quantization (PTQ), especially vector quantization (VQ), focuses on model compression to below 2 bits per parameter, but usually overlooks the data movement bottleneck that dominates real-world inference latency. In this paper, we propose **M**ulti-level **B**lock-wise **V**ector **Q**uantization (MBVQ), a hardware-friendly VQ framework designed to address this issue. MBVQ introduces two key innovations: (1) capacity-constrained per-block codebooks for efficient memory residency during dequantization, and (2) multi-level residual quantization scheme with row-wise importance selection and per-block scaling to compensate for reconstruction error. Specifically, we first partition weight matrices into blocks to adapt large-scale codebooks to varying hardwares. Subsequently, we merge redundant adjacent blocks for quantizing residuals with selective codebooks and design codebook rescaling to adaptively adjust the block-wise distribution of values to simultaneously improve model accuracy and achieve substantial inference speedup. Experiments on a range of LLMs demonstrate that MBVQ consistently outperforms state-of-the-art low-bit quantization baselines under comparable compression ratios, delivering approximately inference speedup, while maintaining competitive perplexity and zero-shot accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.