Restoring Backdoor Detection under Adaptive Attacks in Multimodal Large Language Models
Abstract
Backdoor detection in multimodal large language models (MLLMs) often relies on predefined statistics of hidden representations or attention. Detector-aware attackers can suppress these statistics during fine-tuning while preserving effective backdoors. We investigate whether detection can be restored after such evasion. In the evaluated settings, defeating an existing detection rule leaves poisoned and benign examples distinguishable under a newly learned score. Motivated by this finding, we propose Benign-Anchored Witness Detection (BAWD), which restores detection by relearning a scoring function on the attacked model. BAWD trains lightweight source classifiers to distinguish audited representations from a trusted same-task benign reference. Given this reference, fitting and thresholding require no additional poisoning labels or exact poisoning rate; thresholding uses a fixed candidate-rank window. Under a matched-reference assumption, the population-optimal source score induces the same ordering as the poisoning posterior, without requiring access to it. Evaluations span three MLLMs and three datasets. On Qwen3-VL-8B fine-tuned on ScienceQA, the linear variant (BAWD-L) achieves 99.8–100.0% detection F1 across seven detector-aware attack configurations, while each of seven baselines degrades substantially under at least one configuration. In two four-round attack sequences targeting previously deployed BAWD detectors, freshly fitted BAWD-L achieves 99.5–99.8% F1 while attack success remains at 99.5–99.8%. These results demonstrate that relearning scores with trusted benign references can restore detection after the evaluated detector-aware attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.