acceptodds
Under review as a conference paper at ICLR 2027

Error-Guided Refinement and Composition for LLM Quantization

Abstract

Post-training quantization has emerged as an active research area for enabling the efficient deployment of large language models (LLMs). A wide range of quantization methods have been proposed, each based on different design intuitions. We observe that one method can have smaller quantization errors on some weights, while another has smaller errors on others. These differences can affect model performance and may reflect how each method prioritizes different weights. Motivated by this observation, we develop an error-guided approach for both improving individual quantizers and combining their strengths. We first show that selecting more accurate weight approximations from different methods can improve model quality, but introduces additional storage to record the selected method for each weight. To address this issue, we propose an error-guided conversion method that constructs non-uniform variants of existing quantizers using their quantization errors. This conversion provides a common codebook-based representation for different methods. We then combine their codebooks into a single codebook with the same number of entries as an individual non-uniform quantizer, avoiding additional per-weight method labels. Experiments on multiple LLMs show that the proposed conversion can improve perplexity over the original uniform quantizers, while codebook composition provides further gains in several evaluated settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.