acceptodds
Under review as a conference paper at ICLR 2027

CLEAN FEATURE ANCHORING: MITIGATING THE ACCURACY–ROBUSTNESS TRADE-OFF BY SELF-DISTILLING A NATURALLY TRAINED MODEL

Abstract

Adversarial training improves robustness but often sacrifices accuracy on clean data, leading to a persistent clean–robust trade-off. We address this trade-off by first learning a high-accuracy clean self-teacher from natural images and then preserving its clean knowledge while adversarially training the student. We propose Clean Feature Anchoring (CFA), which matches the student's representation of an adversarial input to the teacher's fixed representation of the corresponding clean input while retaining the teacher's classifier. We show that controlling this fixed clean anchor simultaneously preserves agreement with the teacher on clean inputs and limits representation variation within the perturbation neighborhood. We also find that teacher clean accuracy alone does not determine how well clean performance is retained after robust training, indicating that transferability depends on properties beyond teacher accuracy. Finally, because the same input-space perturbation can induce substantially different changes in the anchoring objective across examples, we introduce sensitivity-matched perturbation budgets that assign smaller radii to more sensitive examples and larger radii to less sensitive ones while preserving the average perturbation budget. Across CIFAR-10, CIFAR-100, and Tiny-ImageNet, CFA consistently combines high clean accuracy with strong adversarial robustness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.