TARE: Weigh a Never-Poisoned Twin Before Reading Backdoor-Defense Costs
Abstract
Backdoor-defense leaderboards print a clean-accuracy drop and read it as the cost of removal. The drop is measured on the poisoned victim alone, so it cannot separate removal from what the defense does to any model, and it inherits where the victim started. Where the victim starts is, for three of BackdoorBench's sixteen attacks, a configuration file: WaNet, BPP and Input-Aware ship a MultiStepLR whose first milestone is the epoch count, so their victims never anneal and are the lowest-accuracy victims in 30/31 public CIFAR cells at ≤5%. On PreAct-ResNet18, fine-tuning-family defenses return a low start to a level of their own, so on those victims the published cost is negative, the benchmark's rating clips the "gain" to zero, and 2 of the 48 citing defense papers we read rest a no-cost claim on those cells; re-run with their released code, TSBD and CGD "gain" on a never-poisoned model too. We isolate the cause with a 2×2 that edits only that scheduler line, its swapped arms self-registered before they ran: the sign of the fine-tuning family's clean-model cost reverses both ways while its published gain on the annealed victim only shrinks toward zero, 44/44 seeds following the schedule, replicated on BPP, FT-SAM, CIFAR-100 and VGG19-BN and induced in a second toolkit. The instrument is TARE: run the identical defense on a never-poisoned twin of the same recipe, schedule and seed (on BackdoorBench, ≤10 poisoned images, admitted only below 5% attack success); what the twin loses is the tare. On BackdoorBench's BadNets grid seven of the eight defenses charge the twin (Neural Cleanse only where its detector fires), +0.13 points (fine-tuning) to +5.70 (I-BAU); the eighth, ABL, destroys it. Within an attack the start cancels from rankings, so the tare re-orders nothing there; what poisoning adds beyond it is printed under two estimators and not corrected, its removal share being unidentified. We ship the three-key patch, a signed tare column over 7 attacks × 8 defenses of the CIFAR-10 roster, and TARE-Z, a twin-free estimator for seed-stable defenses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.