From Triggers to Antigens: An Immune-Inspired Framework for Continual Backdoor Defense
Abstract
Deep neural networks are increasingly deployed in dynamic environments where models must continuously learn from evolving data streams, yet many existing backdoor defenses remain fundamentally limited by their reliance on predefined trigger patterns. Such a trigger-centric view fails when attacks become distributed, representation-dependent, or semantically consistent with their target labels, leaving the underlying malicious behavior invisible to conventional detectors. In this work, we reveal a different perspective: backdoors are not characterized by their observable triggers, but by recurring non-self behaviors that persist after normal semantic responses are removed. Through extensive empirical analysis, we discover that samples carrying a recurring backdoor factor across different semantic contents can share a class-conditioned residual, analogous to a pathogen-derived antigen exposed after subtracting healthy self responses. This observation motivates a new formulation of continual backdoor defense as an antigen-to-antibody learning process rather than a trigger identification problem. Building upon this insight, we propose Immune-inspired Semantic Network (ISN), a continual defense framework that learns transferable immune memory against emerging backdoor behaviors. ISN establishes a stable semantic self-reference, extracts attack-invariant antigens through complementary receptors, and transforms validated antibodies into adaptive immune responses including correction, neutralization, and vaccination. Unlike many existing approaches that memorize individual attacks, ISN accumulates reusable immune memory capable of recognizing recurrent threats across tasks and attack families. Empirical results demonstrate that our framework achieves the state-of-the-art performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.