acceptodds
Under review as a conference paper at ICLR 2027

Rethinking the Quantization Error Lifecycle in LLMs

Abstract

Post-training quantization (PTQ) is essential for efficient deployment of large language models (LLMs), yet aggressive low-bit quantization can severely degrade model accuracy. Existing GPTQ-style compensation-based methods sequentially quantize weight columns, typically determining the quantization decision of the current column using round-to-nearest (RTN) and then updating the remaining weights to compensate the induced error. However, the RTN-based quantization decision is made independently of subsequent compensation, thereby treating error formation and error compensation as separate stages. In this work, we uncover the overlooked coupling between error formation and error compensation and formalize this coupling through the quantization error lifecycle, a systematic perspective on compensation-based LLM quantization. Building on this perspective, we derive a shared compensation representation that reveals complementary information for error formation and error compensation. From this shared representation, we develop CARE-Q, a unified compensation-based quantization framework comprising DiCR for compensation-aware rounding and ODPP for enhanced error compensation. Extensive experiments across diverse LLMs and quantization settings show that CARE-Q consistently outperforms leading compensation-based quantization methods and complements representative transformation-based approaches, including QuaRot and FlatQuant, while introducing no additional inference-time overhead. Our code is provided in the supplementary material and at https://anonymous.4open.science/r/CARE-Q/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.