Prediction Instability and Task Utility in Low-Precision Language Models
Abstract
Selective precision and quantization repair are evaluated through changes in predictions, reconstruction error, and task accuracy. We examine how these measurements relate using a controlled audit of low-precision language models. A prespecified experiment evaluates two instruction-tuned models on 1,024 HellaSwag training examples across four precision paths and three execution layouts. Changing option batch size while preserving padding changes per-item precision-upgrade utility on 42 examples for Qwen and 14 for SmolLM2; corresponding full-model FP32 arithmetic controls show zero changes. Within fixed 256-example margin selections, opposing changes cancel, leaving aggregate shifts of -6 and +1 net-correct units. Subsequent SmolLM2 repair diagnostics compare four teacher-target objectives with the same rank-16 parameterization. Centered-score fitting reduces training error by 46.17%, while error on two evaluation panels increases by 15.24% and 8.16%. On the previously observed 1,024-question transfer panel, it yields three additional correct answers, compared with five for option-level distillation. An ARC replay further distinguishes offline fallback utility from executed predictions and cost. The measurements identify where prediction stability and reconstruction fidelity diverge from task utility, and provide a protocol for evaluating these quantities together.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.