Adaptive Mixed-Precision Codebook Quantization for Large Language Models
Abstract
Codebook quantization enables efficient compression of large language models to low bit-widths, making it suitable for resource-constrained deployment. Existing methods assign a uniform codebook configuration to all matrices. However, under a tight average-bit budget, this uniform configuration leads to capacity misallocation: insensitive linear maps receive disproportionate precision, while sensit ive projection layers remain under-represented. To address this issue, we propose Adaptive Mixed-Precision Codebook Quantization (AmpCQ), a framework that assigns different codebook configurations based on per-matrix sensitivity. The proposed method combines: (i) a heterogeneous tier space comprising single- and dual-codebook quantizers; (ii) a normalized activation-aware reconstruction metric for cross-matrix comparability; (iii) Pareto-filtered greedy allocation guided by marginal return per bit; and (iv) lightweight codebook-delta adaptation for downstream tasks. On Llama3-8B and Qwen3-14B, AmpCQ demonstrates superior performance over uniform AQLM-style allocation at matched average bit width. Specifically, when Qwen3-14B is compressed to an average bit width of approximately 3 bits, it retains 97.3% of its original average zero-shot accuracy. Our code is available in the Supplementary Material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.