RampQ: Quantizing What Recovery Cannot Fix for Mixed-Precision LLM Quantization
Abstract
Low-bit quantization reduces the memory footprint of large language models but often degrades model quality at aggressive precision. Mixed-precision quantization alleviates this problem by assigning different bit-widths based on module sensitivity or quantization error, while low-rank recovery adds a low-rank branch to recover the quantization error. Pre-recovery quantization error, however, does not necessarily reflect the error that remains after recovery. We show that quantization-error recoverability varies substantially across modules and bit-widths, such that recovery can reorder where higher precision is most beneficial. This motivates allocating precision according to the error that survives recovery. We propose RampQ, a Recovery-Aware Mixed-Precision Quantization framework that performs allocation in the recovered representation space. RampQ constructs bit-specific representations, evaluates each precision choice using module sensitivity and post-recovery residual error, and selects representations under their joint memory cost. A model-level refinement further searches complete-model assignments using selective end-to-end perplexity feedback. Our evaluation results show that RampQ improves the quality memory trade off over mixed precision and low rank recovery baselines. Compared with the best baseline under comparable memory settings, RampQ reduces perplexity on average by 8.1% on WikiText-2 and 5.7% on C4, and improves average zero-shot accuracy by 1.67 percentage points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.