acceptodds
Under review as a conference paper at ICLR 2027

Teacher–Student Disagreement For Uncertainty-Aware Human-In-The-Loop AI

Abstract

Large-scale classifiers are vulnerable to adversarial test-time attacks: small, deliberately crafted input perturbations that cause confident misclassification. Defenses such as adversarial training give only empirical robustness, and formal certification does not scale to modern models, so a deployed system still needs a mechanism for deciding when a prediction should be deferred to a human expert. Such selective prediction schemes typically use the model's own confidence, its maximum softmax probability, as the deferral trigger. However, our analysis shows that under attack this score is not merely unreliable but inverted: on CIFAR-10 and MNIST, attacked inputs receive higher confidence than clean ones, so a confidence-based rule defers less exactly when it should defer more. We propose a framework that pairs a standard-trained teacher model with a small ensemble of lightweight student models that are distilled from the teacher, adversarially trained, and uses teacher–student disagreement as a signal that the teacher may have been compromised. Against predictive entropy, mutual information and vote margin computed on the ensemble, and against nine selective-prediction and adversarial-detection baselines, disagreement is the strongest signal for deciding which teacher predictions to trust, that is, for ranking the teacher's wrong predictions ahead of its correct ones so that a human reviews the wrong ones. For instance, it ranks teacher errors with an area under the ROC curve (AUROC) of – where confidence scores –. Adversarial detectors such as Mahalanobis distance detect attacks as well as disagreement does but cannot rank teacher errors. Because attacks crafted for the teacher rarely transfer to the robust students, the ensemble also serves as a fallback predictor (– accuracy where the teacher is at –). The same mechanism sets the limits: on clean inputs confidence remains the better selector; measures computed on the students alone are blind to attacks that do not transfer to them; the students must be accurate enough to serve as a reference, a condition the practitioner can check from clean accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.