acceptodds
Under review as a conference paper at ICLR 2027

FTTrap: Downstream Fine-Tuning Walks into a Backdoor Trap

Abstract

Fine-tuning is widely used to adapt LLMs to downstream tasks. However, prior work has shown that fine-tuning on benign data can cause unintended behavioral changes, such as increased vulnerability to jailbreaks. The deliberate exploitation of such vulnerabilities remains largely underexplored. Specifically, an attacker could implant a backdoor that remains hidden at release and reactivates after fine-tuning. Achieving this is challenging, as different fine-tuning methods and datasets produce vastly different parameter changes, making consistent backdoor activation difficult. To solve this issue, we propose FTTrap, a dormant backdoor activated by downstream fine-tuning. FTTrap first trains the model to produce backdoor responses to target inputs. It then restores normal responses without erasing the learned backdoor behavior. The restoration is deliberately fragile, allowing downstream fine-tuning to activate the backdoor. We evaluate FTTrap on 6 models (2B–27B), 4 attack scenarios, and 12 downstream tasks, including 4 related to the corresponding attack scenarios and 8 unrelated. At release, ASR remains at most 1.0% across all settings. After fine-tuning on 8 unrelated datasets, the mean ASR rises to 80.2% with minimal utility degradation. FTTrap demonstrates strong generalization across diverse adaptation settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.