PoisonTrace: Tracing Data-Poisoning Triggers through Input-Hessian Shifts
Abstract
A trigger-based dirty-label backdoor attack implants a trigger that forces a text classifier to output an attacker-chosen label while leaving its behavior on clean inputs intact. We consider a defender who holds a clean pre-trained checkpoint, its fine-tuned counterpart and the poisoned fine-tuning dataset, and who must identify the trigger tokens without any knowledge of the trigger or the poisoning configuration. For this setting, we introduce PoisonTrace, an end-to-end trigger-localization framework based on per-probe input-Hessian analysis. PoisonTrace first infers the trigger-associated class and selects a small set of suspicious training examples as probes, using prediction shifts between the two checkpoints. Its core step then estimates the input-Hessian block-row energy at every probe position with randomized Hessian–vector products, and keeps the positions whose within-probe share of curvature grows after fine-tuning. Aggregating this evidence by tokenizer identity yields a ranking of suspicious tokens that is robust to trigger placement. Finally, a score-gap rule estimates how many top-ranked tokens to flag as the trigger. We evaluate PoisonTrace on 26 triggers derived from NIST TrojAI Round 5, ranging from single symbols to multi-sentence passages, inserted at random word boundaries in four corpora: three sentiment datasets and WildGuardMix(WGM), a safety-moderation dataset for large language model guardrails. PoisonTrace exacts trigger-token set with the F1 scores of 78.3% to 94.1%. Scrubbing the identified tokens from the training data and retraining lowers the attack success rate by 49.6–79.8, while clean accuracy changes by less than one percentage point on three of the four corpora.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.