acceptodds
Under review as a conference paper at ICLR 2027

Sensitivity-Guided Mixed-Precision Quantization for Large Language Models

Abstract

Post-training quantization (PTQ) is essential for deploying large language models (LLMs) on resource-constrained hardware, but its impact on specific capabilities is poorly characterized. We present a capacity-controlled study of how weight-only quantization affects multi-step reasoning versus factual recall and commonsense inference, across six models (0.5B–8B; Qwen2.5, Qwen3, and Phi-3.5 families). We introduce the Reasoning Fragility Ratio (RFR), a diagnostic metric that quantifies capability-asymmetric degradation, and SGMPQ, a mixed-precision scheme that assigns 8-bit precision to the layers ranked most quantization-sensitive by a perturbation probe, with a capability-conditioned ratio as the ranking signal. Our evaluation uses real benchmarks (GSM8K, ARC-Challenge, HellaSwag, held-out perplexity) with Wilson confidence intervals and exact McNemar tests, and compares SGMPQ 8/4 (≈5.2 average bits) against capacity-matched controls: uniform W5/W6, a random mixed-precision control at the identical budget (replicated over five allocation seeds), and—decisively—two capability-agnostic sensitivity rankings at the same budget, plus GPTQ and AWQ at matched budgets. Three findings emerge. First, the quantizer dominates where it matters: GPTQ beats RTN by up to 63 accuracy points exactly where RTN breaks (W4 on the 3B base and Coder-3B models), rescuing Qwen2.5-Coder-3B, which RTN destroys outright, while the two quantizers converge on robust models at W5. Second, on robust models the bit budget masks placement: at equal ≈5.2 bits, sensitivity-guided placement is statistically indistinguishable from random placement on the two small instruct models, worse on the 3B base model, and better on the two largest models—though with five allocation seeds the permutation design (sign-flip floor p=0.0625) cannot reach conventional significance even if the advantage is real on every seed. Third, placement matters exactly where uniform quantization fails: on Coder-3B, both perturbation-probe-guided allocations tested preserve the model (59.4–64.2% GSM8K) where all five independently drawn random allocations collapse (0.4–2.4%; uniform W5 also rescues it), and the capability ratio itself is beneficial on one model (8B: +6.0 pp) but actively harmful on another (3B: -33.4 pp vs. the agnostic ranking). We release per-item results, sensitivity maps, and the full evaluation stack, and conclude that sensitivity-guided mixed precision is a targeted repair tool for quantization-brittle models—but no ranking signal we tested (capability ratio or agnostic) transfers safely across models, so any allocation must be validated against random-placement controls on the target model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.