acceptodds
Under review as a conference paper at ICLR 2027

Beyond 4 Bits: Data-Centric Post-Training Quantization and Recovery

Abstract

Post-training quantization (PTQ) aims at compressing the weights of a large trained model using a limited compute and data budget. Down to 4 bits, state-of-the-art PTQ methods are near-lossless; below this limit, the tested scalar PTQ configurations collapse to near-random accuracy and model quality becomes acutely sensitive to choices such as the quantization representation and the calibration data. At such low bitwidths, recovery fine-tuning after quantization becomes crucial, and its effectiveness depends heavily on these same choices. Yet there is little systematic evidence on how the quantization method, the quality and task alignment of calibration data, and the design of the recovery procedure interact to determine the best final 2-bit model. We present a systematic study of 2-bit quantization and recovery fine-tuning across multiple model families, scales, and data distributions, and distill our findings into a practical, architecture-dependent set of best practices for efficiently obtaining an accurate 2-bit model. This recipe, which combines a vector codebook representation, a small task-aligned calibration mixture, and a short codebook-only distillation stage, restores %–% average accuracy across seven question-answering tasks at 2 bits (a %–% bf16 ceiling), versus the % random accuracy floor of scalar PTQ (more so on MoE than dense models), at roughly half the compute budget of the strongest Quantization Aware Training (QAT) baseline. Our contribution is empirical: a systematic, architecture-aware map of the 2-bit design space, together with released code, Hessian mixtures, and checkpoints that make the findings directly actionable and reproducible.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.