acceptodds
Under review as a conference paper at ICLR 2027

RampQ: Quantizing What Recovery Cannot Fix for Mixed-Precision LLM Quantization

Abstract

Low-bit quantization reduces the memory footprint of large language models but often degrades model quality at aggressive precision. Mixed-precision quantization alleviates this problem by assigning different bit-widths based on module sensitivity or quantization error, while low-rank recovery adds a low-rank branch to recover the quantization error. Pre-recovery quantization error, however, does not necessarily reflect the error that remains after recovery. We show that quantization-error recoverability varies substantially across modules and bit-widths, such that recovery can reorder where higher precision is most beneficial. This motivates allocating precision according to the error that survives recovery. We propose RampQ, a Recovery-Aware Mixed-Precision Quantization framework that performs allocation in the recovered representation space. RampQ constructs bit-specific representations, evaluates each precision choice using module sensitivity and post-recovery residual error, and selects representations under their joint memory cost. A model-level refinement further searches complete-model assignments using selective end-to-end perplexity feedback. Our evaluation results show that RampQ improves the quality memory trade off over mixed precision and low rank recovery baselines. Compared with the best baseline under comparable memory settings, RampQ reduces perplexity on average by 8.1% on WikiText-2 and 5.7% on C4, and improves average zero-shot accuracy by 1.67 percentage points.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.