ReSC-Q: Residual-Aware Reliable PTQ with Shrinkage-Controlled Compensation
Abstract
Post-training quantization (PTQ) reduces the memory footprint and inference cost of LLMs under tight deployment constraints. Hessian-based PTQ methods such as GPTQ use estimated curvature information for error compensation, but compensation estimated from calibration data does not always lead to optimal reconstruction loss on unseen inputs. This becomes especially problematic in ultra low-bit quantization. We propose ReSC-Q, which controls uncertain cross-group curvature through shrinkage while explicitly accounting for the reconstruction loss remaining after compensation. ReSC-Q supports practical group-wise quantization by separating the blockwise reconstruction objective into cross-group compensation and within-group quantization, applying shrinkage and residual-aware weighting to each, respectively. Across seven LLMs evaluated at 2-bit, ReSC-Q improves average zero-shot accuracy by 1.09 percentage points over the strongest baseline for each model, reducing the mean accuracy gap to full precision by 23.2%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.