The Harmful Tail: Toward Lossless Quantization of Large Language Models
Abstract
Low-bit quantization substantially reduces the cost of LLM inference, yet the relationship between error magnitude and predictive degradation remains poorly characterized. We show that this relationship is highly non-uniform across spectral directions. In W4A4-quantized Qwen3 models, dominant high-energy components of the quantization error contribute relatively little to performance loss whereas a low-energy spectral tail explains most of the excess loss. A gold aware analysis further reveals that correcting only four input-conditioned tail directions removes 0.43% of tail-error energy, while increasing mean downstream accuracy beyond the BF16 reference. We attribute this energy function mismatch to loss geometry. Tail directions exhibit systematically higher directional curvature per unit error energy. Motivated by this, we introduce Harmful-Tail-Guided Sparse Compensation, HTSC, which identifies modules whose local corrections propagate beneficially into harmful tail directions and distills this repair into sparse low-rank branches with a frozen backbone. At inference, the reference model and spectral decomposition are not required. Across Qwen3 models, HTSC restores performance to near or slightly above the BF16 reference, with low additional overhead. Code is available at: https://anonymous.4open.science/r/HTSCptq-CE19/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.