When Quantization Preserves Generation but Degrades Verification: A Frozen-Candidate Study across Model Scale
Abstract
Quantized language models are commonly evaluated as generators, yet the same models are increasingly used to judge, rerank, and select candidate answers. These roles need not respond identically to reduced precision. We study this distinc- tion with a frozen-candidate protocol: for each model, eight candidates are gen- erated once in FP16 and then held byte-identical while the verifier is evaluated at FP16 and several quantized precisions. Correctness is assigned by determinis- tic semantic checking, so verifier comparisons are not confounded by a changing candidate pool. On 300 MATH-500 problems, Qwen2.5-7B-Instruct provides a clear example. HQQ 8-bit changes greedy generation accuracy from 0.467 to 0.473 (∆ = +0.006), while global verifier AUROC falls from 0.825 to 0.544 and within-problem ranking from 0.698 to 0.411. Best-of-eight selection falls from 0.500 to 0.430, below the exact random-selection rate of 0.455. The same 7B degradation is observed with bitsandbytes int8 and NF4. In contrast, the HQQ 8- bit result does not show a comparable verification drop at Qwen2.5-3B, -14B, or -32B; the 3B result is additionally sensitive to the quantizer, with NF4 producing a separate degradation. A Llama-3.1-8B experiment is not a strong cross-family test because its FP16 verifier is already near chance (AUROC = 0.471). Output- level analysis of the 7B model shows an approximately 8×reduction in correct- versus-incorrect margin separation and a higher decision-flip rate among correct candidates than incorrect candidates. The evidence therefore supports a narrower claim than a universal “quantization breaks verification” hypothesis: generation accuracy alone does not certify preservation of verifier ranking, and verification should be evaluated directly for the model, precision, interface, and candidate set used downstream.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.