ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
Abstract
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods typically initialize quantized models with heuristic rules or greedy optimization, and usually treat the resulting integer assignments as final. We introduce ReQuant, a backpropagation-free fixed-grid refinement stage *within* PTQ that optimizes an executable quantized model while preserving its format. Agnostic to the PTQ initializer, ReQuant starts from an existing feasible quantized model and iteratively revisits its discrete weight assignments on the fixed quantization grid. Accepted updates strictly reduce the mean squared reconstruction error and remain on the original grid. ReQuant thus integrates with existing PTQ pipelines as a plug-and-play post-processing stage. Experiments across diverse model families, bit-widths, and downstream tasks show that ReQuant consistently improves quantized models from heterogeneous PTQ initializers, with especially large gains on simple initializers and lower bit-widths. Notably, ReQuant can refine a simple round-to-nearest initialization across multiple sweeps until it approaches or surpasses GPTAQ under the same quantization format.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.