Enhancing Robustness through Repair-Oriented Honeypots and Selective Knowledge Aggregation
Abstract
A model with established adversarial robustness can remain vulnerable to an encountered attack. Targeted improvement requires useful repair data and a learning procedure that limits damage to existing capabilities. We introduce HoneyTeacher, a general framework that couples repair-oriented honeypot induction with selective knowledge aggregation. The production model's estimated repair benefit guides honeypot configuration, shaping attack interactions to yield training material with greater downstream learning value. A material teacher learns from these data; aggregation then selects supervision from this teacher and the original model according to their joint correctness, consolidating complementary corrections while retaining existing strengths. We evaluate score-based black-box and gradient-based white-box instantiations on all 10,000 CIFAR-10 test images. Under Square and PGD attacks, respectively, target robust accuracy improves by 12.4% and 8.7% relative to the original model, and by 17.1% and 13.5% relative to ordinary-material direct learning. On the secondary Square evaluation, the white-box model improves accuracy by 24.0% relative to the original model and 34.2% relative to ordinary-material direct learning. Clean accuracy changes by +0.17 and −1.26 percentage points relative to the original model in the black-box and white-box settings, respectively, with all monitored capabilities meeting the adopted retention criteria. Induced material improves robustness under matched direct learning, and aggregation adds further gains; together, they reverse the net target-robustness losses observed with ordinary direct learning in both settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.