acceptodds
Under review as a conference paper at ICLR 2027

Accuracy Is Not Enough: Measuring Instruction-Following Stability Under Post-Training Quantization

Abstract

Post-training quantization is routinely evaluated by task accuracy alone, leaving open whether the instruction-following behavior instilled during instruction tuning survives aggressive compression. We study this question for INT4 weight-only quantization of two compact instruction-tuned models, Llama 3 8B Instruct and Phi-3 Mini 3.8B Instruct, across seven prompt structures and three decoding temperatures on GSM8K and the MMLU high-school mathematics subset (12,600 generations). We score every generation both for task accuracy and for strict compliance with an explicit output contract, and we show that the parse-rate measure used in prior practice is saturated above 0.99 and cannot detect the effects we report. Three findings emerge. First, task correctness and instruction-level compliance decouple under INT4, and the compliance effect is the larger one: Phi-3 contract compliance falls from 0.107 to 0.055 while its accuracy simultaneously rises from 0.344 to 0.373, so a purely accuracy-based evaluation would record this model as unaffected or improved. Second, the effect is strongly model-dependent: Llama 3 compliance is stable and marginally better under INT4, and the two models also re-rank prompt effectiveness differently on accuracy. Third, the compliance loss is largest in exactly the exemplar-based prompts that appear protective on accuracy, so few-shot prompting cannot be assumed to stabilize behavior under compression. We discuss a distributional hypothesis that is consistent with these behaviors but that we do not test directly, and we argue that evaluation of quantized models should treat instruction-following stability as a first-class metric alongside accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.