acceptodds
Under review as a conference paper at ICLR 2027

FidelityQ: Preserving Predictive Fidelity in Ultra-Low-Bit Language Model Quantization

Abstract

Ultra-low-bit scalar quantization offers a regular arithmetic structure for efficient deployment of large language models, but preserving their predictive behavior at extreme precision remains challenging. With only a few reconstruction levels, accurate weight representation must translate into consistent model predictions using a small set of calibration samples. We propose FidelityQ, a scalar quantization framework for preserving predictive fidelity at weight precisions down to 2 bits. Specifically, we introduce Information-Aligned Analytic Coding (IAC) for accurate quantization. IAC shapes each weight row toward a centered symmetric source through randomized Hadamard mixing, then fits a mirrored four-level minimum-distortion quantizer to its empirical weights. Every level set in this symmetry-aligned family decomposes exactly into 2 sign-valued bases with associated scales, combining source-adaptive reconstruction with lookup-free decoding compatible with scaled integer arithmetic. Our analysis further reveals that residual predictive discrepancies are heavy-tailed and concentrate on tokens that the teacher predicts confidently, motivating Knowledge-Intensive Calibration (KIC) for consistent quantization. KIC ranks sequences using the tail of token-wise teacher-student divergence and allocates a limited calibration budget to the highest-scoring samples. It then restores predictive agreement by updating only quantization levels and scales under teacher supervision, leaving the pretrained weights frozen. Experiments on the LLaMA and Qwen families demonstrate the effectiveness of FidelityQ across weight-only and joint quantization settings. On LLaMA-2-7B, FidelityQ achieves a WikiText2 perplexity of 6.62 at W2A16 and improves the prior state-of-the-art W2A4KV4 result from 8.31 to 6.96, while reducing linear-layer weight memory by 8.0× relative to FP16. These results support accurate LLM deployment with structured ultra-low-bit quantization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.