acceptodds
Under review as a conference paper at ICLR 2027

Patronus: Identifying and Mitigating Transferable Backdoors in Pre-trained Language Models

Abstract

Transferable backdoors pose a potentially high-impact supply-chain risk to Pre-trained Language Models (PLMs): a compromised checkpoint may retain attacker-chosen behavior after downstream adaptation. Several defenses work on continuous representation signatures, but signatures extracted from the original checkpoint may shift when model parameters changes in downstream fine-tuning. To address this, We present Patronus, an input-centric framework that directly searches a suspicious PLM for discrete candidate triggers. Patronus formulates trigger recovery as source-supervised multi-trigger contrastive learning and performs gradient-guided discrete search over candidate triggers. The recovered candidates are then screened on held-out inputs before mitigation. Verified candidates support exact-match input filtering and adversarial-training-based model purification. In model-level detection, Patronus detects 435 of 436 valid backdoored instances (99.77%). For exact trigger recovery, Patronus achieves higher ground-truth-trigger recall with lower search cost than a representative trigger-inversion baseline. For downstream mitigation, Patronus outperforms the evaluated defense baselines in attack success rate (ASR) across nine downstream datasets, reducing ASR toward clean-model levels while preserving downstream task performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.