ImmuneCLIP: Adaptation-Stable Defense Against Backdoor Rebound in Purified CLIP Models
Abstract
Multimodal contrastive learning models such as CLIP align images and text in a shared representation and are widely adapted to downstream tasks. Yet existing backdoor purification defenses are typically evaluated only immediately after purification, leaving their security during later adaptation largely unexplored. We identify backdoor rebound, a failure mode in which a purified model initially achieves a low attack success rate (ASR) but regains strong backdoor behavior after benign downstream fine-tuning. By analyzing fine-tuning update directions, we find that clean optimization can move the model along parameter-space directions that reactivate residual backdoor behavior. To address this threat, we introduce ImmuneCLIP, an adaptation-stable defense that protects purified models beyond the initial purification checkpoint. ImmuneCLIP first uses trigger inversion to construct a differentiable proxy for residual backdoor risk. It then collects parameter-update directions from representative clean adaptation procedures to build an adaptation probe bank. Based on these probes, we define reactivation susceptibility as the largest first-order increase in proxy risk along clean adaptation directions. ImmuneCLIP jointly minimizes residual backdoor risk and reactivation susceptibility at the purified checkpoint and at nearby states reached through simulated clean updates. Contrastive regularization and knowledge distillation are further used to preserve clean performance. We evaluate four backdoor attacks, four purification defenses, and five adaptation pipelines. ImmuneCLIP's worst-case post-adaptation ASR anchors range from 2.0±0.7% to 9.5±1.0%, with final clean-accuracy anchors between 50.1±0.4% and 54.9±0.4%. The evaluation measures both delivery-time suppression and resistance to backdoor rebound during subsequent adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.