DISTRIQUANT: BEYOND RECONSTRUCTION FOR ROBUST BINARY LLMS
Abstract
Post-training quantization (PTQ) for large language models (LLMs) is commonly formulated as a reconstruction problem, where quantized weights are optimized to reproduce full-precision behavior on a small calibration set. We show that, in the 1-bit regime, reconstruction plays a more consequential role: it implicitly determines which representation directions survive quantization. Since reconstruction error places greater emphasis on directions with high variance on the calibration distribution, binary PTQ preferentially preserves calibration-dominant structure while weak directions receive substantially less protection. Crucially, directions that are weak on the calibration distribution are not necessarily unimportant elsewhere. We find that many become relatively more prominent on unseen domains, yet are disproportionately suppressed after binarization. To quantify this effect, we introduce a directional survivor rate that measures how much of each full-precision representation direction remains after quantization. Motivated by this finding, we propose DistriQuant, a spectrum-balanced approach to salient-mask selection for binary PTQ. DistriQuant uses Covariance-Balanced Alignment to remove the variance imbalance among calibration directions before comparing full-precision and quantized representations, preventing dominant calibration directions from overwhelming the allocation of limited binary capacity. The method requires only the standard calibration set, leaves the underlying binary representation unchanged, and introduces no additional inference-time computation. Extensive experiments across different architectures and datasets demonstrate the effectiveness of our proposed method.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.