SPEAR: Piercing Stealthy Backdoors in Vision-Language Models with Energy Signatures
Abstract
VLMs have transformed multimodal understanding, yet they remain vulnerable to stealthy backdoor attacks that traditional input-centric or distillation-based defenses often fail to detect. We observe that although the triggers may be visually subtle, successful trigger-target shortcuts consistently induce measurable shifts in the model's internal attention-energy profiles, which we formalize as Heterogeneous Energy Divergence (HED). Based on this discovery, we propose a data purification framework called SPEAR (Surrogate-based Poisoned-sample Energy Analysis and Removal). SPEAR utilizes Energy-Driven Surrogate Learning (EDSL) to transform latent triggers into measurable energy anomalies and a Dual-Path Perception (DPP) mechanism to excise poisoned samples by intersecting global distribution shifts with local layer-wise spikes. Extensive evaluations across multiple benchmarks and VLM architectures show that SPEAR reduces ASR to near-zero while in most cases improving clean reasoning fidelity under equal-volume tuning. Additional stress tests under ultra-low poisoning rates, diverse VLM-effective triggers, existing VLM backdoor attacks, and an adaptive energy-regularized attacker suggest that SPEAR captures a recurring internal energy signature in evaluated backdoor shortcuts rather than a fixed trigger artifact. The code is available at https://anonymous.4open.science/r/backdoorvlm-071D/link.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.