From Hidden States to Probability Measures: Measuring Quantization-Induced Deformation in Large Language Models
Abstract
Attention rows and next-token output distributions are probability measures. Prior work evaluates compressed models with accuracy, output KL, and answer flips; we extend this output-level evaluation to attention routing and treat both distributions as a measure-valued diagnostic layer. We compare discrepancies with and without ground geometry to characterize quantization-induced redistribution over positions and vocabulary outcomes. Using same-prefix, teacher-forced comparisons isolates local deformation from free-running history divergence. In the tested, comparatively mild INT8-versus-FP8 weights-only comparison, attention KL and \(W_1\) distinguish all six model-domain pairs with a consistent direction; perplexity is significant in only \(3/6\) and disagrees in direction even among significant pairs. Output KL has one directional exception. Under stronger INT4 weights-only quantization, perplexity also detects degradation. Adding activation quantization reverses the INT8/FP8 ordering in every tested pair, with model- and diagnostic-dependent magnitude and architecture-dependent layerwise shape. These diagnostics measure local redistribution and are not downstream-accuracy surrogates or full path-law comparisons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.