acceptodds
Under review as a conference paper at ICLR 2027

When Poisoning Fails: Reassessing Backdoor Risk in Object Detection

Abstract

Backdoor attacks pose a severe threat to deep learning, yet their behaviour in object detection remains understudied. While attacks have been proposed for object detection, we revisit how they are evaluated and find that they are less effective than initially thought. In particular, Attack Success Rate (ASR) for Region Misclassification Attacks (RMA) does not reflect cases where the original class is retained alongside the target class, Mean Average Precision (mAP) is an indirect proxy for Object Disappearance Attacks (ODA), and trigger scaling relative to the object size matters. When these factors are accounted for, existing data-poisoning attacks are less effective than their original evaluations suggest, and even at 100% poisoning fail to consistently achieve reliable backdoor behaviour across detector architectures. Motivated by these findings, we study the outsourced training threat model standard in image-classification backdoor research. Within this setting, we introduce BadDet+, a unified penalty-based framework that augments the detector loss with a log-barrier term to selectively suppress confident original-class predictions on trigger-bearing objects. Across COCO, MTSD, and physical-world PTSD benchmarks spanning four detector families, BadDet+ provides reliable backdoor behaviour while preserving clean-task performance, demonstrating that object detectors remain vulnerable under this alternative threat model. This finding has important implications for future defensive efforts and the practical deployment of object detectors.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.