SPARSE: An Explainability-Inspired Defense Against Neural Trojans
Abstract
Deep Neural Networks are increasingly available as pretrained checkpoints from public hubs or fine-tuned by third-party providers. This exposes them to Trojan attacks: an adversary poisons the training data and/or implants a backdoor during training so that the CNN emits an attacker-chosen output on inputs containing a specific trigger. Existing defenses either target a single architecture or, when representation-based and architecture-agnostic, operate offline on a poisoned training set rather than flagging individual test inputs online. None has been validated across CNN, encoder-only, and decoder-only LLM architectures within a single evaluation framework. We propose SPARSE (SAE-based Poisoned Anomaly Recognition and Selective Exclusion), an online defense method built on a hypothesis coming from mechanistic interpretability: a trigger is an out-of-distribution signal that the backdoor keys on, so it should activate a distinct subset of the model's internal "concepts" that the clean inputs almost never activate. SPARSE attaches a small SAE to a chosen layer of the target model as a read-only side-car (leaving the forward pass unmodified), fits it together with a lightweight anomaly detector over its fit statistics computed using only clean data, and flags any input whose SAE output departs from the output of clean-data concepts. The same mechanism works for CNNs, encoder-only LMs and generative LLMs. Across 4 models, 12 datasets, and 14 attacks, SPARSE achieves an F1-score of 87%, outperforming the best baselines by 15 percentage points with a common SAE-based detection mechanism, instantiated per modality, for CNNs, encoder-only LMs and generative LLMs. The implementation of SPARSE is publicly available at https://anonymous.4open.science/r/SPARSE-54FD
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.