CostBit: Candidate-Cost-Informed Bit Allocation for Post-Training Quantization of LLMs
Abstract
Post-training quantization (PTQ) is a practical route to reducing the memory and bandwidth requirements of large language models (LLMs). However, uniform bit-width quantization becomes increasingly inefficient in the aggressive sub-3-bit regime due to heterogeneous quantization sensitivity across weight blocks. Although existing mixed-precision schemes alleviate this issue, their allocation rules fail to account for the quantization error variation of each weight block across candidate bit‑widths, resulting in suboptimal accuracy–compression trade-offs. To address this limitation, we introduce CostBit, a candidate-Cost-informed blockwise Bit allocation method for PTQ of LLMs. Specifically, CostBit evaluates the Hessian-weighted cost associated with each candidate bit-width for every weight block and leverages a Lagrangian allocator to closely match the target bit budget, supporting non‑integer values. Extensive experiments on OPT, LLaMA-2, and LLaMA-3 spanning 125M to 70B parameters demonstrate that CostBit yields improved performance at comparable target bit budgets, with significant gains in the sub-3-bit regime. Further, at an average bit-width of approximately 2.5 bits, it reduces WikiText-2 perplexity from 7.69 to 6.43 and improves the average zero-shot accuracy from 60.04% to 63.79% compared with OWQ on LLaMA-2-7B. CostBit also enables fully GPU-resident inference of LLaMA-2-70B on a single RTX 4090 (24 GB) at approximately 2.5 bits, achieving 75.70% average zero-shot accuracy, only 1.37% below the FP16 baseline. Our code is available at https://anonymous.4open.science/r/CostBit.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.