Adaptive Decision-Boundary Probing for Robust Distillation
Abstract
Robust distillation has been shown to effectively transfer robustness from larger adversarially trained models to smaller student models. However, robust teacher models may not reliably classify all adversarial examples, and conventional distillation mostly do not explicitly account for the position of these examples relative to the teacher's decision boundary. In this work, we propose Adaptive Decision-Boundary Probing Distillation (ADBD), a robust distillation framework that uses the teacher's classification behavior to construct sample-specific query points. Starting from an adversarial example generated through coordinated teacher-student information, we preserve its perturbation direction while adaptively scaling its magnitude. When the teacher misclassifies the adversarial example, the perturbation magnitude is reduced to obtain a query point closer to the teacher's decision boundary while remaining on the teacher-correct side. When the teacher correctly classifies the adversarial example, the perturbation is extended along the same direction to probe a farther region in which the teacher remains correct. The resulting query samples provide teacher supervision at locations selected according to the teacher's local decision-boundary behavior. We evaluate ADBD across multiple datasets and adversarial attacks. Our experimental results show that ADBD improves the adversarial robustness of distilled student models over existing robust distillation methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.