Data-Chain Backdoor: Do You Trust Diffusion Models as Generative Data Supplier?
Abstract
The increasing use of generative models such as diffusion models for synthetic data augmentation has greatly reduced the cost of data collection and labeling in downstream perception tasks. However, this new data source paradigm may introduce important security concerns. Publicly available generative models are often reused without verification, raising a fundamental question of their safety and trustworthiness. We present Data-Chain Backdoor (DCB), a framework that registers existing backdoor triggers in diffusion models for reproduction in clean-label synthetic data. DCB generates images containing backdoor triggers from ordinary target-class prompts without assuming specific user wording, enabling passive backdoor propagation through synthetic data augmentation. We establish this threat through generator training from scratch and develop an efficient fine-tuning pipeline for trigger registration. Guided by our analysis of trigger registration and reproduction, we further introduce ChainTint, a trigger designed for faithful propagation through DCB and stealth against backdoor defenses. Extensive experiments across diffusion models, downstream classifiers, and defense settings evaluate DCB with existing triggers and ChainTint, including applications to data-scarce learning and dermatological image classification. The DCB framework and ChainTint demonstrate a practical threat to downstream models in generative data supply chains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.