RIDE: A Lightweight Plug-and-Play Framework for Enhancing Adversarial Robustness
Abstract
Deep neural networks (DNNs) remain vulnerable to adversarial attacks based on deliberately crafted perturbations. Previous arts target specific adversarial attacks (e.g., adversarial training) or impose substantial computational overhead on inference time (e.g., adversarial purification). Moreover, both of them significantly compromise the accuracy on clean examples, which dominate real-world tasks, thereby marginalizing their practical applications. To address these challenges, we propose a lightweight plug-and-play framework, namely RIDE, to protect DNNs in inference time. By combining detection and purification, RIDE precisely identifies and purifies adversarial examples, enhancing model robustness while maintaining comparable accuracy on clean examples. In particular, RIDE leverages information from the low-variance subspace (a subspace that exhibits severe selection bias) to precisely detect adversarial examples, which have been demonstrated to diverge from clean examples in this subspace. Moreover, the samples identified as adversarial are purified through an adversarial filter, which uses a denoising autoencoder to optimize the input toward regions of higher training data coverage, driving adversarial examples back to the clean data manifold. Finally, RIDE applies inference-time data augmentation to reduce noise in the predictions, further enhancing the robustness of this framework. Extensive experiments demonstrate that RIDE increases robust accuracy at most by 50.97% while maintaining comparable accuracy on clean examples. Moreover, RIDE uses only 0.87% of the inference time compared to that of classical inference-time defense methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.