acceptodds
Under review as a conference paper at ICLR 2027

Theoretically Scaled Quantization: A Data-Free Analytical Framework for LLMs

Abstract

Large Language Models (LLMs) require efficient post-training quantization (PTQ) to reduce memory and computational costs while preserving model quality. Existing PTQ methods commonly rely on calibration data, activation statistics, Hessian information, or iterative optimization, introducing additional preprocessing overhead and scale storage. We propose Theoretically Scaled Quantization (TSQ), a fully data-free analytical PTQ framework that derives quantization parameters directly from pretrained weight statistics. We formulate the TSQ Distortion Objective, , an exact distribution-independent geometric distortion objective for finite multi-region quantization, and define the optimal geometric scaling factor through one-dimensional minimization over . We further derive exact region-wise absolute-error bounds that decrease geometrically with region index and a region-independent relative-error bound of for all non-innermost regions. Under the empirically observed near-Gaussian structure of pretrained LLM weights, we derive the TSQ Governing Optimality Equation, , which analytically characterizes the optimal scaling factor using the normalized dynamic range , eliminating calibration data, activation statistics, Hessian computation, and iterative optimization. Building on the regional distortion analysis, we derive a generalized TSQ distortion law that decomposes normalized distortion into the weight dynamic range, optimized geometric distortion, and quantization resolution. TSQ also employs a scale-efficient representation requiring only one per-row maximum and a single global scaling parameter, from which all thresholds and quantization step sizes are analytically derived, reducing scale storage from to . Experiments across OPT, Llama, and Qwen model families show that TSQ achieves competitive or lower perplexity than AWQ across the evaluated settings while requiring no calibration data, substantially less scale storage, and lower full-model quantization time. These results demonstrate that analytical geometric scaling can provide a data-free and scalable approach to LLM post-training quantization.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.