One Token Owns the Calibration Hessian: Why Layer-wise Quantization Can Silently Break Instruction Following
Abstract
Quantized instruction-tuned models are usually validated with perplexity and multiple-choice accuracy, which can miss a collapse of instruction following. At 3 bits, reference GPTQ with act-order, calibrated on short documents as widely used libraries sample them by default, lowers the IFEval score of Llama-3.1-8B-Instruct from 0.768 to 0.155 and that of Qwen2.5-14B-Instruct from 0.820 to 0.426, below round-to-nearest (0.565, 0.697), while MMLU stays in the ordinary 3-bit range; GPTQModel's own column order avoids the collapse. We trace the failure to a blind spot of the layer-wise objective. At the down-projection that forms the attention sink, the first token owns the calibration Hessian, yet the loss is almost insensitive to its output. GPTQ's compensation, which we derive in closed form, therefore protects that token and writes its rounding error along the sink pattern into one matrix: a lesion that the chat-template tokens read out. Excising the lesion restores both models. The blind spot is present in all twenty-six models whose first token dominates a down-projection, and of the seventeen of these that we quantized, fourteen carry the lesion, yet only two collapse. Transplants indicate that the other matrices of a healthy network cancel the lesion, which would explain why the failure is rare and depends on how the calibration text is cut and which documents are drawn. Weighting each calibration token by its output-side sensitivity divided by its input energy removes the blind spot and the four collapses it was run on, at a cost of at most 5.6 points on healthy models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.