KLIQ: KL-Informed Low-Bit Quantization via Attention-Distribution-Aware Fisher-Kronecker Curvature
Abstract
Post-training quantization (PTQ) reduces the memory and computation costs of large language models (LLMs), but preserving accuracy at very low bit-widths remains challenging. Existing reconstruction-based curvature approximations in second-order PTQ can overlook the distribution-dependent sensitivity of attention probabilities to query and key quantization errors. We introduce KLIQ, a backpropagation-free PTQ method that uses KL divergence between full-precision and quantized attention distributions to guide curvature estimation, quantization parameter selection, and error compensation. KLIQ introduces three key innovations: (1) a Fisher-Kronecker curvature surrogate that maps a diagonal approximation of the attention-score categorical Fisher to query and key weight spaces; (2) staged functional calibration that conditions key quantization parameter selection on already quantized and compensated queries to account for query-key error interactions; and (3) KL-derived compensation statistics for compensating quantization errors propagated from preceding layers. Extensive experiments across multiple model families show that, when combined with conventional outlier-suppression techniques, KLIQ achieves state-of-the-art performance on language modeling and downstream tasks under low-bit weight-only and weight-activation quantization. The source code is available at https://anonymous.4open.science/r/KLIQ-7440/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.