CAJQUANT: LARGE LANGUAGE MODEL QUANTIZATION VIA COMPENSATION-AWARE JOINT GRID SELECTION
Abstract
Large language models (LLMs) impose substantial memory and inference costs, motivating weight-only post-training quantization (PTQ) for efficient deployment. However, preserving low-bit accuracy while keeping quantization overhead low remains challenging. In column-wise quantization with error compensation, the candidate grid determines the quantization errors in the current column, which drive compensation updates to weights in later columns. Static group-wise grid evaluation overlooks these candidate-dependent updates by scoring grids on the weights before intra-group compensation. Moreover, fixing the zero point during clipping search can overlook grids with smaller quantization errors for weights outside the clipping interval. To address these issues, we propose Compensation-Aware Joint Grid Selection for LLM Quantization (CAJQuant), which jointly selects the clipping ratio and zero point using compensation-aware grid evaluation. This evaluation measures reconstruction loss under optimal compensation while accounting for the weight updates induced by each candidate grid. Experiments on Llama models from 7B to 70B parameters demonstrate strong performance at 2, 3, and 4 bits, with substantial improvements in model quality at 2 bits. In this setting, CAJQuant outperforms all evaluated baselines in WikiText-2 perplexity and average downstream accuracy across all four models. On a single NVIDIA A100-80GB GPU, CAJQuant completes 2-bit quantization of Llama-2-70B with a group size of 128 in approximately 5 hours.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.