acceptodds
Under review as a conference paper at ICLR 2027

RCDQ: Reference-Guided LLM Quantization with Interleaved Continuous Compensation and Discrete Decisions

Abstract

Full-precision reference outputs are widely used to compensate for accumulated errors in low-bit quantization of large language models. Existing methods incorporate these references in different ways. Some apply continuous compensation to an entire layer before quantization. Others incorporate deviations from reference outputs into sequential weight updates or use reference-based reconstruction objectives to optimize quantization parameters and integer codes. A key challenge is that improvements in output reconstruction achieved through continuous compensation may not be preserved after discretization. To address this challenge, we propose RCDQ (Reference-Guided Continuous–Discrete Quantization), a two-stage group-wise quantization method that uses the same full-precision reference objective to guide both continuous compensation and quantization decisions. In stage one, before quantizing each weight group, we recompute its weight compensation using statistics describing the output discrepancy between the quantized and full-precision paths. These statistics are updated as groups are processed sequentially. Under the same full-precision reference, in stage two, we then evaluate fake-quantized candidates constructed from the compensated weights and select the best quantization scale multipliers to reduce the discrepancy between the candidate outputs and the full-precision reference outputs. Experiments across seven Llama and Mistral models show that RCDQ achieves the lowest W3 perplexity in all 14 WikiText-2 and C4 model–dataset settings, together with the highest or tied-highest six-task average and MMLU accuracy on six of seven models, respectively. RCDQ remains robust at more aggressive precisions, achieving the lowest W2A16 perplexity in 10 of 14 settings and the lowest W3A8/W4A8 perplexity in three of four settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.