Not All Reconstruction Errors Matter: Score-Guided Block-wise Reconstruction for Low-Bit LLM Quantization
Abstract
Block-wise reconstruction has been widely used for post-training quantization (PTQ) of large language models (LLMs), but conventional objectives penalize all block-output perturbation directions equally despite their different effects on the final predictive distribution. In this paper, we propose SABRE, a score-guided anisotropic reconstruction approach that accounts for these unequal predictive effects during block-wise reconstruction. To identify prediction-sensitive directions, SABRE approximates the Fisher curvature of the full-precision (FP) predictive distribution using rank-one outer products of sampled score vectors. Based on this approximation, SABRE assigns additional penalties to errors aligned with these directions while retaining isotropic reconstruction to control the overall perturbation magnitude. To ensure stable weighting across blocks and tokens, we further introduce gradient-norm normalization that bounds token-dependent directional weights, thereby reducing sensitivity to variations in score-vector scale. The FP-derived score vectors are computed once and cached, enabling efficient block-local reconstruction without repeated end-to-end evaluation. Because SABRE modifies only the reconstruction objective, it can be readily integrated into existing block-wise quantization methods without changing their optimization variables or inference-time computation. Extensive experiments across diverse quantization methods, LLM architectures, and low-bit settings demonstrate consistent perplexity improvements, substantial recovery under aggressive 2-bit quantization, and broad downstream gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.