acceptodds
Under review as a conference paper at ICLR 2027

FROM ATTENTION SENSITIVITY TO LAYER ROLE: REVISITING MIXED-PRECISION QUANTIZATION OF TRANSFORMERS

Abstract

Most post-training quantization pipelines keep each quantized weight matrix close to its pretrained counterpart, one matrix at a time, and stop there. Whether that proxy tracks what the attention block actually computes, or how errors in the Q, K and V projections compound once they meet inside the softmax, is rarely checked. The obvious fix is to write the objective on the attention output itself, over all three projections at once, and to use that same objective wherever the pipeline needs a signal. That is what JAB does. It defines one scalar loss over the joint Q, K, V weights of a block, evaluated against the block's real causally-masked attention output, and uses it both to fit the quantized weights (GPTQ warm start, then STE with learnable scales) and to score the block for a multiple-choice knapsack allocation. On attention-only quantization of Mistral-7B this pays off: at 3 bits JAB recovers – of the gap between uniform GPTQ and full precision, and its sensitivity estimate tracks an oracle costing 73 forward passes to within a fraction of a perplexity point. Once MLP layers enter the allocation, though, JAB is beaten by a role-aware offset rule that needs no sensitivity estimate at all, both on GPT-2 with QKV and MLP quantized together and on the full Mistral-7B model. With a 3-bit floor, this rule quantizes of Mistral-7B's weights to 4.5 bits per parameter at perplexity — within of full precision () at compression, 14 GB to 3.93 GB — against for JAB at the same budget, and to 3.5 bits at (). Whichever matrix a weight belongs to inside its block matters more than any sensitivity estimate we managed to compute. Two things came out sideways. Block-local reconstruction turned out to be an unreliable proxy for end-to-end perplexity: in one run a improvement in a block's own objective came with a increase in perplexity, which is why every allocation in this paper is validated end-to-end rather than on local objectives alone. And on attention-only quantization, fine-tuning consistently moved weights farther from their pretrained values while pulling attention outputs closer, with net gains. Post-training seems to recover attention behavior, not weights.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.