acceptodds
Under review as a conference paper at ICLR 2027

Interpretable Adversarial Defense via Class-Aware Neuron Sensitivity Attribution and Recalibration

Abstract

Research on analyzing and enhancing model robustness at the neuron level has attracted significant attention due to its interpretability. However, existing methods are either based on neuron sensitivity evaluated by changes in activation values, or based on neuron importance measured by contributions to clean samples. Neither approach considers the intrinsic connection between the changes in neuron activations and the neural network's outputs. To bridge this gap, we propose a novel Neuron Sensitivity Attribution (NSA) method. By jointly calculating the dual contributions of intermediate neurons in suppressing the true label and promoting the adversarial label, this method provides a fine-grained, category-specific evaluation of neuron sensitivity under adversarial perturbations. Based on this metric, we propose a Sensitivity-Aware Neuron Recalibration mechanism to reform adversarial training strategies. Specifically, during adversarial training, stronger penalties are assigned to highly sensitive neurons, thereby enabling the model to rely more heavily on robust neurons. Extensive experiments across various models and adversarial training frameworks demonstrate that the proposed method can serve as an effective plug-and-play module to further enhance the robustness of deep learning models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.