Quantization Can Preserve Verification While Degrading Self-Generated Supervision
Abstract
Can a quantized model retain its ability to select answers yet lose the learning value of its own solutions? We separate verification ranking, final-answer correctness and held-out learning gain. On Gemma 4 12B and Qwen3.6-35B-A3B, all nine registered 4- and 3-bit verifier contrasts are equivalent to full precision within ±0.025 AUROC, even as generation degrades. Crossing sample source with learner precision reveals that SmolLM3-3B's full-precision solutions at 38% correctness teach both learners more than NF4 solutions at 74% (by 1.9 and 3.6 accuracy points) on a holdout selected for partial solvability. On the 12B, approximately matching sample count and problem coverage leaves an 18.9-point source advantage for the quantized learner on this holdout and 4.6 points on MATH-500. Completed, mostly correct NF4 solutions also reduce the full-precision learner's accuracy. The penalty increases along the GGUF precision ladder on both learning subjects, with format-dependent severity, and no tested selector shows an established recovery. Reciprocal continuations separately localize the 12B's excess truncation to the continuing model for the tested prefixes. Under this math supervised fine-tuning recipe, verification robustness and final-answer correctness are insufficient summaries of useful self-generated supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.