JOLTing Robust Aggregation with Trigger Optimized Clean-Image Backdoor Poisoning
Abstract
In this work, we study the attack surface of planting backdoors with clean-image poisoning in models defended by robust aggregation. This setting is relevant because labels can be easier to control than inputs (e.g., likes/dislikes in social media or crowd-sourcing), and cause potentially stealthier attacks than feature manipulation. However, clean-image poisoning offers more limited control over the training process, making backdoor implantation more challenging. First, we provide empirical evidence that FLIP, a state-of-the-art gradient-based clean-image backdoor attack, fails under robust aggregation, across datasets, and architectures. Second, we propose the Joint Optimization of Labels and Trigger (JOLT) method. We show that, in both centralized and federated settings, and without knowledge of the defense, JOLT successfully plants backdoors with high attack success rates, while maintaining a good clean test accuracy and small poisoning budget. Along the way, we identify an inconsistency between the implementation of FLIP and the described method, which gives a misleading impression of efficiency. To the best of our knowledge, JOLT is the first clean-image poisoning attack that successfully plants backdoors under robust aggregation defenses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.