When Stronger Triggers Backfire: A High-Dimensional Theory of Backdoor Attacks
Abstract
Backdoor poisoning attacks behave counter-intuitively in high dimensions: when the test-time trigger budget is held fixed, a stronger training trigger can help the defender. We study generalised linear models on Gaussian-mixture data in the proportional regime (), varying the training trigger strength . Three phenomena emerge: (i) clean test accuracy can increase with ; (ii) the trigger alignment that drives attack success peaks at a finite and then declines; and (iii) low-variance directions can yield stronger trigger alignment. We derive closed-form asymptotic alignments for squared loss, including a unique finite peak for arbitrary mean-trigger overlap, and extend the finite-peak result and large-strength decay to a broader class of convex losses under suitable assumptions. A margin-variance decomposition identifies a finite-sample term that helps explain (i) and is absent from the population limit. Experiments on Gaussian surrogates and logistic regression on CIFAR-10 support the theoretical predictions, while ResNet-18 experiments exhibit similar qualitative behaviour beyond the convex setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.