Interpretable backdoor detection using sparse concept representations
Abstract
Inference-time backdoor detection methods aim to identify model inputs that contain a hidden trigger. However, few explain why an input is flagged as poisoned. Without such explanations, the detector's decisions are difficult to audit, contest, and learn from. To address this, we propose VARDE (Verdict Attribution via Representation DEcomposition), a backdoor detection method that uses human-interpretable concepts. VARDE transforms the activations into such concepts using a sparse autoencoder at every layer. The verdict is based on the Mahalanobis distance between present concepts and the estimated class-conditional distribution of concepts. The distance decomposes additively over concepts, providing an exact value of how much a particular concept at a specific layer contributed to the verdict. For CNNs, this decomposition maps onto the model's input space; each concept's contribution splits exactly across spatial locations, yielding a concept map that localises it in the input. VARDE is computationally efficient, as the single forward pass producing the model prediction also yields the verdict and its explanation. We show for a variety of image classification datasets how VARDE explains its verdict while being competitive with state-of-the-art inference-time detectors. We also compare VARDE against Grad-CAM, commonly used in backdoor detection. VARDE's concept maps are more faithful than Grad-CAM's saliency maps, consistently concentrating more strongly on the trigger region.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.