Distill Globally, Train Locally: Error-Localized Recovery for 4-bit LLMs
Abstract
Low-bit quantization enables efficient large language model (LLM) inference, but maintaining accuracy with 4-bit weights and activations (W4A4) remains challenging. Local calibration in post-training quantization (PTQ) lacks direct supervision of final predictions, limiting further accuracy recovery. Full-model quantization-aware distillation (QAD) can recover near-BF16 accuracy through global supervision, but requires costly model-wide optimization. To preserve the accuracy of full-model QAD while minimizing training cost, we propose LoDiQ, a lightweight error-localized distillation framework with three components: (1) propagated-error localization identifies sensitive input channels; (2) quantized residual compensation strengthens recovery within expanded low-bit GEMMs, with Gradient-Guided Residual Compensation Allocation (GRCA) supporting lower-budget configurations; and (3) localized distillation trains only the selected corrective weights using compact teacher targets, avoiding full-vocabulary teacher-logit storage. Experiments across the Qwen and Llama model families span language modeling, zero-shot commonsense reasoning, broad-domain knowledge, and mathematical reasoning. LoDiQ surpasses state-of-the-art PTQ baselines in downstream accuracy under both MXFP4 and NVFP4 W4A4 quantization. On a matched reasoning evaluation, LoDiQ outperforms full-model QAD using only 1/100 of the distillation tokens and 4.67% of its trainable parameters, substantially reducing optimization workload and trainable-state memory. Inference benchmarks on RTX 5090 show up to 3.20× and 3.14× time-to-first-token (TTFT) speedups over BF16 in MXFP4 and NVFP4, respectively.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.