Projected-Distortion Guided Quantization
Abstract
Post-training quantization (PTQ) reduces the memory and latency costs of large language models (LLMs), but existing methods degrade sharply at 2–3 bits, in part because codebook fitting, outlier handling, and pruning are designed and tuned separately. We present Projected-Distortion Guided Quantization (PDGQ), which drives all three with a single activation-aware measure of output distortion. PDGQ combines (1) a maximum-entropy codebook refined by activation-weighted Lloyd iterations, (2) Projected-Distortion Outlier Restoration (PDOR), which restores weights whose distortion is large relative to a rate–distortion reference, on a scale shared across layers and bit-widths, and (3) Distortion-Consistent Pruning (DCP), which zeroes weights that zero represents better than their codeword. On LLaMA and Qwen3 models, PDGQ matches recent PTQ methods in perplexity and zero-shot accuracy at 4 bits and outperforms them by a margin that grows as the bit budget shrinks, being largest at nominal 2–3 bits, with a few narrow exceptions that we report. With simple activation and KV-cache quantization, PDGQ also preserves DeepSeek-R1-Distill accuracy on GSM8K and MATH-500 under W4A4KV4 better than FlatQuant and QuaRot.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.