Contextual Utility of Quantization Moves in Extreme Low-Bit LLMs
Abstract
Accurate measurements of isolated weight replacements need not predict their joint effect on a quantized model. We study this gap for legal replacements at a fixed bitrate, separating isolated pricing error, composition error and changes in evaluation data. On Llama-3.2-1B and activation-ordered 3B GPTQ bases, option-KL composition residuals are 26× and 42× the sums of absolute midpoint-pricing errors; large residuals remain with exact-singleton prices. Exhaustive Qwen3-0.6B lattices reveal context-dependent effects: six of eight moves in the stratified lattice have statistically resolved ARC-KL sign reversals, excluding an exact monotone additive representation. Pair measurements improve prediction of unqueried combinations on both Qwen and Llama. Their value for selection depends on the criterion: the stratified Qwen front retains a quality-coverage advantage after refitting, whereas the 3B additive and quadratic fronts are nearly equal in quality at a one-standard-error tolerance. Low-order endpoint structure leaves continuous gradient responses undetermined, and shared-midpoint prices can differ from those of the executed subset. Support-matched controls show that restricting the number of replacements curbs the failures of large combinations, while additional endpoint search gives bank-dependent NLL/ARC trade-offs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.