acceptodds
Under review as a conference paper at ICLR 2027

ARSO: Reconstruction-Aware Rounding with Joint Scale Optimization for Weight-Only LLM Quantization

Abstract

Post-training quantization (PTQ) compresses large language models (LLMs) into low-precision representations to reduce memory and accelerate inference without retraining. In weight-only PTQ, existing methods either rely on round-to-nearest (RTN), whose local decisions ignore reconstruction error, or employ costly iterative rounding optimization. Moreover, quantization scales are often reused for dequantization, coupling integer-grid construction with weight reconstruction. We identify reconstruction freedom as an important optimization resource. Scale Optimization (SO) exploits scale freedom by deriving reconstruction-optimal dequantization scales for fixed integer codes; eliminating this scale freedom further reveals a directional criterion over integer weights, enabling Reconstruction-Aware Rounding (AR). In sequential quantization, AR exploits future compensation freedom through a conditional reconstruction metric, while J-SO exploits prefix scale freedom by jointly updating processed-group scales. This stage-wise exploitation naturally yields ARSO, a PTQ method with single-pass reconstruction-aware rounding that achieves the lowest perplexity in 32 of 36 settings and the highest average zero-shot accuracy across all six models at 2 bits among the evaluated methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.