acceptodds
Under review as a conference paper at ICLR 2027

Clean Data, Dirty Secrets: Backdoor Transfer in LLM Distillation via Defect Amplification

Abstract

Knowledge distillation (KD) is widely used to distill knowledge from large language models (LLMs) to compact student models, enabling efficient deployment in resource-constrained scenarios. It is commonly considered that using clean distillation data can mitigate the transfer of explicit trigger-based backdoors from poisoned teacher models to student models. In this paper, we propose DefAttack, a novel backdoor attack that enables backdoor transfer even when the student is distilled on clean data. Unlike existing explicit trigger-based backdoors that rely on transferable trigger-response associations, DefAttack exploits latent model defects that manifest as systematic biases in the teacher model's output logits. Specifically, we introduce a defect amplification loss during teacher poisoning, which increases the logits assigned to target-domain tokens while constraining changes to benign output logits. This transforms latent defects into transferable backdoors, allowing student models to implicitly inherit them while aligning with teacher supervision on clean distillation data. Importantly, DefAttack decouples backdoor transfer from activation: the defect backdoor is inherited during clean-data distillation and activated only afterwards using an optimized suffix. Evaluations across four models and three tasks demonstrate an average attack success rate of at least 95.3% and a false positive rate of at most 4.5%. Further analyses show that defect backdoors rely primarily on transferred defects rather than the optimized suffix.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.