acceptodds
Under review as a conference paper at ICLR 2027

Distilling Answers, Discarding Uncertainty: Uncertainty Collapse in Sequence-Level Knowledge Distillation

Abstract

Sequence-level knowledge distillation (SeqKD) commonly trains a student on a single sampled teacher response per input. While this transfers the teacher’s answer, it need not transfer the uncertainty with which that answer was produced. We study the resulting uncertainty collapse and its consequences for hallucination detection. In expectation, the SeqKD objective is minimised by matching the teacher’s response distribution and therefore preserves its uncertainty. But with one fixed sampled target per input, the empirical objective instead favours concentration on that target, regardless of whether it is correct. We show theoretically that a single response cannot identify the teacher’s dispersion over alternative meanings, and that training provides a systematic pressure towards increasingly concentrated student output distributions. Across four teacher–student pairs on SimpleQA, continued distillation causes students to increasingly reproduce hallucinated teacher responses while becoming more certain of them: semantic-entropy AUROC for detecting incorrect answers falls from roughly 0.8 to chance, even when answer accuracy changes little. The effect also appears, more weakly, on held-out and out-of-distribution data, and recurs in mathematical reasoning. We evaluate filtering, abstention replacement, and multi-response distillation as mitigations. Among the interventions we study, retaining multiple teacher responses directly preserves distributional information and substantially maintains hallucination detectability, at a cost that grows with the number of responses. These results show that transferring a teacher’s outputs and transferring its uncertainty are distinct objectives, and that reliable distillation requires preserving more than a single draw from the teacher’s response distribution.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.