acceptodds
Under review as a conference paper at ICLR 2027

Randomized Feature Squeezing against Unseen Attacks without Adversarial Training

Abstract

Deep learning has made tremendous progress in the last decades; however, it is not robust to adversarial attacks. The most effective approach is perhaps adversarial training, although it is impractical because it requires prior knowledge about the attackers and incurs high computational costs. In this paper, we propose a novel approach that can train a robust network only through standard training with clean images without awareness of the attacker's strategy. We add a specially designed network input layer, which accomplishes a randomized feature squeezing to reduce the malicious perturbation. It achieves excellent robustness against unseen and attacks at one time in terms of the computational cost of the attacker versus the defender through just 100 and 50 epochs of standard training with clean images in CIFAR-10 and ImageNet respectively. The thorough experimental results validate the high performance. Moreover, it can also defend against unlearnable examples generated by One-Pixel Shortcut which breaks down the adversarial training approach.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.