Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models
Abstract
Transferable backdoors pose a potentially high-impact supply-chain risk to Pre-trained Language Models (PLMs): a compromised checkpoint may retain attacker-chosen behavior after downstream adaptation. Several defenses work on continuous representation signatures, but signatures extracted from the original checkpoint may shift when model parameters changes in downstream fine-tuning. To address this, We present Patronus, an input-centric framework that directly searches a suspicious PLM for discrete candidate triggers. Patronus formulates trigger recovery as source-supervised multi-trigger contrastive learning and performs gradient-guided discrete search over candidate triggers. The recovered candidates are then screened on held-out inputs before mitigation. Verified candidates support exact-match input filtering and adversarial-training-based model purification. In model-level detection, Patronus detects 435 of 436 valid backdoored instances (99.77%). For exact trigger recovery, Patronus achieves higher ground-truth-trigger recall with lower search cost than a representative trigger-inversion baseline. For downstream mitigation, Patronus outperforms the evaluated defense baselines in attack success rate (ASR) across nine downstream datasets, reducing ASR toward clean-model levels while preserving downstream task performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.