RC-VQ: Recovering What Calibration Discards with Vector-Quantized Residual Correction
Abstract
Calibration data decides what training-free LLM compression keeps. In activationaware factorization and post-training quantization alike, a small calibration set defines an importance measure over the weights, and whatever it does not certify is discarded as a residual. We argue that failing the calibration proxy does not make weights unimportant, and that the discarded residual should be recovered. The standard recovery tool, a low-rank adapter, is structurally mismatched to this task. On LLaMA-2 7B the residuals are dense and spectrally flat. Quantization residuals are statistically indistinguishable from a Marchenko–Pastur noise bulk, and truncation residuals form a gapless spectral continuum without isolated spikes, so a rank-r corrector recovers a vanishing fraction of either. Across four model families, the energy a low-rank corrector recovers grows only linearly in its rank. We propose RC-VQ, a full-rank vector-quantized corrector fitted to the residual without gradients or recovery data, which composes with SVD factors, INT4 weights, and their stacked combination. On structural compression of LLaMA-2 7B, LLaMA-3 8B, and Qwen3 8B, RC-VQ reduces perplexity degradation by 3–13× against the strongest reported structured baseline at matched checkpoint size. On the INT4-AWQ quantization gap of Qwen3.5 4B, it recovers 96% of the distance to FP16 at three bits per parameter, a budget at which an identically sized low-rank corrector recovers under 40% and a scalar residual branch needs 37% more bits to match. Sweeping codebook objectives across structural, quantization, and stacked residuals, we show that each pure objective has a failure regime, while the hierarchical direction-then-refinement objective, PolarVQ, never collapses, making it the robust default when residual geometry is unknown.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.