Distilling Bayesian Uncertainty into a Single Forward Pass
Abstract
Deep ensembles give reliable Bayesian uncertainty estimates, but they need forward passes for every input. We treat the uncertainty of a Bayesian teacher as a function of the input and train a single deterministic network, our student, to output it in one forward pass. Distilling such posterior quantities has been done before, but the targets are Monte Carlo estimates from members: they change when the ensemble is retrained, and what a student learns from these noisy targets has not been analysed. We model the student's uncertainty head as a regression on the noisy teacher uncertainty, allowing correlated noise and a biased teacher, and give an exact condition under which the student is closer to the ideal uncertainty than its teacher. When the teacher's bias and the part of its noise that does not average out over members are small, the best regularisation strength decreases as grows. A controlled simulation and ensembles trained from scratch follow the predicted trend in . To learn the uncertainty function, our student model first learns the teacher's predictions; a small head on the frozen backbone then learns the teacher's uncertainty on clean, corrupted, mixed and masked inputs, with a log-space regression loss and a ranking loss. Any teacher that provides this uncertainty can be used. On image classification, segmentation and text, our student stays close to the teacher's accuracy at the cost of one forward pass. Scored against ground-truth labels on image classification, it has the lowest or tied-lowest selective-prediction risk among single-pass methods that do not distil the teacher, is competitive with other distillation methods, and detects distribution shift without labels as well as the teacher. Its out-of-distribution detection is close to the teacher's.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.