acceptodds
Under review as a conference paper at ICLR 2027

Budgets Transfer, Hardness Labels Do Not: Auditing What Survives Quantization of Reasoning Models

Abstract

Reasoning models are usually quantized just before deployment, but two settings that guide their use are measured earlier, on the full-precision (BF16) model: the decode budget, which caps how many tokens the model may generate, and prompt hardness labels, which decide which prompts receive extra training. We ask whether these BF16 measurements can be reused after quantization. A direct comparison is misleading, because a longer budget can look better only because fewer answers are cut off, and because two runs of the same model already disagree on which prompts are hard. We therefore check truncation before comparing budgets, and compare every cross-precision change in the hard set with an independent rerun of the same model. On Qwen3.5-4B with 4-bit GPTQ, the two settings behave differently. The budget can be reused: serving the 4-bit model at the budget chosen on BF16 changes GPQA accuracy by −0.59 points on average over six seeds (95% interval [−1.94, +0.76]). Hardness labels cannot. The BF16 and 4-bit hard sets overlap by 0.414 (Jaccard), well below the 0.593 between two BF16 runs. The lower overall accuracy of the 4-bit model does not explain the gap: shifting and rescaling the BF16 scores to match it still predicts an overlap of 0.523. Quantization changes which prompts are hard, not only how many. A second quantizer, round-to-nearest, shows the same mismatch. The practical rule is to inherit the budget from BF16 but remeasure hardness on the deployed model, and to test any hardness-based training against a random control, since in our runs it gives no detectable gain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.