acceptodds
Under review as a conference paper at ICLR 2027

Beyond a Single Attack: Generalizable Backdoor Detection via Single-Attack Meta-Classification

Abstract

Backdoor attacks pose a serious security threat to deep neural networks, motivating model-level detection methods that identify whether a target model contains a hidden backdoor before deployment. Meta-classifier-based approaches are particularly attractive because they learn reusable detectors from model-level representations extracted from benign and backdoored reference models. However, we find that existing representations generalize poorly across heterogeneous attacks. Under cross-attack evaluation on CIFAR-10 with ResNet-18, the existing work MNTD and WeightMM achieve average AUCs of only \(0.479\) and \(0.729\), respectively. Our analysis suggests that these representations tend to capture attack-specific characteristics rather than model changes shared across different backdoor attacks. To address this limitation, we propose a novel model-level backdoor detection method by leveraging class-wise decision-space probing. Our key insight is that although backdoor attacks may differ substantially in their triggers and injection mechanisms, a successful attack must ultimately alter the model's decision behavior. We therefore isolate the final classifier and, for each output class, optimize a classifier-input representation that maximizes the corresponding class response. The resulting class-wise representations are aggregated into a model-level representation for meta-classifier training. We evaluate our method under a strict cross-attack protocol covering six dataset–architecture configurations and 12 representative backdoor attacks. Across the evaluated settings, our method achieves an average cross-attack AUC of 0.966, with the best performance reaching 0.998.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.