Consistency Calibration: Post-hoc Confidence from Prediction Stability
Abstract
Calibration is critical for deploying deep learning models in safety-critical applications where reliable confidence estimates guide downstream decisions. Existing studies predominantly view calibration from the reliability perspective, aligning predicted confidence with empirical accuracy, while neglecting whether model predictions remain consistent under perturbations. In this work, we propose Consistency Calibration (CC), a post-hoc method that complements reliability-based calibration by measuring prediction stability. Prior works have explored consistency mainly at the model level, through ensembles or dropout sampling, which aggregate multiple perturbed outputs to smooth predictions but incur substantial computational overhead. In contrast, CC interprets the perturbation process itself as an uncertainty signal, replacing confidence with the empirical stability of predictions under data-, feature-, or logit-level perturbations. We further establish a theoretical connection showing that, under Gumbel logit perturbations, the consistency vector recovers a temperature-scaled softmax distribution and is linked to the Bregman Information of the logit distribution through Renyi entropy. This clarifies the relationship between CC and Temperature Scaling while providing a unified perturbation-consistency interpretation. Extensive experiments on CIFAR-10/100, ImageNet, and ImageNet-LT demonstrate that CC achieves competitive calibration performance compared to both post-hoc and training-time methods, while remaining simple and efficient.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.