Clean Data, Dirty Secrets: Backdoor Transfer in LLM Distillation via Defect Amplification
Abstract
Knowledge distillation (KD) is widely used to distill knowledge from large language models (LLMs) to compact student models, enabling efficient deployment in resource-constrained scenarios. It is commonly considered that using clean distillation data can mitigate the transfer of explicit trigger-based backdoors from poisoned teacher models to student models. In this paper, we propose DefAttack, a novel backdoor attack that enables backdoor transfer even when the student is distilled on clean data. Unlike existing explicit trigger-based backdoors that rely on transferable trigger-response associations, DefAttack exploits latent model defects that manifest as systematic biases in the teacher model's output logits. Specifically, we introduce a defect amplification loss during teacher poisoning, which increases the logits assigned to target-domain tokens while constraining changes to benign output logits. This transforms latent defects into transferable backdoors, allowing student models to implicitly inherit them while aligning with teacher supervision on clean distillation data. Importantly, DefAttack decouples backdoor transfer from activation: the defect backdoor is inherited during clean-data distillation and activated only afterwards using an optimized suffix. Evaluations across four models and three tasks demonstrate an average attack success rate of at least 95.3% and a false positive rate of at most 4.5%. Further analyses show that defect backdoors rely primarily on transferred defects rather than the optimized suffix.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.