acceptodds
Under review as a conference paper at ICLR 2027

The Misallocation of Precision in Quantized Language Models

Abstract

Mixed-precision quantization must decide where to spend a fixed budget of full-precision weights, and the criteria in use are derived from activation statistics: GPTQ minimises layer-wise error under the Hessian , ResQ keeps that covariance's highest-variance subspace at 8 bits. We test the premise underneath, that numerical sensitivity tracks functional importance, by taking its natural per-layer summary, , as a cross-layer allocation rule. In Llama-3.1-8B it is not merely uncorrelated with functional importance but *inverted*: Spearman (, ). The sub-layers it ranks highest sit in a functional dead zone; those it ranks lowest are load-bearing. The cost is measurable: at an identical 6/32 FP16 budget and identical attention/MLP composition, allocating by Hessian rank recovers less of the damage 4-bit quantization does to ARC-Challenge ( vs. , ), and leaves the model significantly worse than FP16 where the functionally-chosen budget does not. We establish this by treating 4-bit NF4 quantization as a controlled perturbation probe, analogous to lesion studies in neuroscience: quantize globally, restore full precision to one sub-layer, measure task recovery. A 64-sub-layer sweep localises the effect to *named sub-layers* rather than to components as classes. Against a do-nothing null measured on the same examples (ARC exactly 0/468) with BH-FDR correction, L13–14 attention clear the null on GSM8K and MATH but on neither measure of ARC, while L6–7 MLP clear it on all three; the component-level contrast is itself not significant (). Split-half selection shows the ranking beats a random equal-budget subset on ARC (pp, 95% CI [, ]) while carrying 4.6–12.2pp of selection optimism. The inversion replicates on Mistral-7B and under GPTQ, and the Llama-3.1-8B-derived layer set transferred to Mistral-7B unchanged recovers of its ARC damage against for Mistral-7B's own Hessian-top layers, though offset by collateral damage elsewhere, so it transfers as a diagnostic rather than an allocator. A surgical model keeping 6 of 32 sub-layers at full precision reaches FP16 equivalence at 42% memory, yet neither Hessian sensitivity nor weight-space SQNR (, ns) predicts which six.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.