FTTrap: Downstream Fine-Tuning Walks into a Backdoor Trap
Abstract
Fine-tuning is widely used to adapt LLMs to downstream tasks. However, prior work has shown that fine-tuning on benign data can cause unintended behavioral changes, such as increased vulnerability to jailbreaks. The deliberate exploitation of such vulnerabilities remains largely underexplored. Specifically, an attacker could implant a backdoor that remains hidden at release and reactivates after fine-tuning. Achieving this is challenging, as different fine-tuning methods and datasets produce vastly different parameter changes, making consistent backdoor activation difficult. To solve this issue, we propose FTTrap, a dormant backdoor activated by downstream fine-tuning. FTTrap first trains the model to produce backdoor responses to target inputs. It then restores normal responses without erasing the learned backdoor behavior. The restoration is deliberately fragile, allowing downstream fine-tuning to activate the backdoor. We evaluate FTTrap on 6 models (2B–27B), 4 attack scenarios, and 12 downstream tasks, including 4 related to the corresponding attack scenarios and 8 unrelated. At release, ASR remains at most 1.0% across all settings. After fine-tuning on 8 unrelated datasets, the mean ASR rises to 80.2% with minimal utility degradation. FTTrap demonstrates strong generalization across diverse adaptation settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.