LLM-Assisted Clean-Label Data Poisoning of Language Models
Abstract
Fine-tuning on externally sourced data exposes language models to data poisoning: an attacker manipulates training examples to influence the learned model. Clean-label backdoor attacks preserve the original labels while inducing targeted predictions on inputs carrying an attacker-chosen trigger. We introduce LLM-CT, an LLM-assisted method that automates dataset-specific candidate generation and combines sentence screening with low-confidence host selection. Shared prompts guide an LLM to infer corpus-specific generation fields and carrier topics, producing reusable sentences for lexical and clean-reference behavioral screening. An accepted sentence is appended to selected target-class hosts, preserving their text apart from normalization. Construction requires no victim queries, weights, or gradients. We report attack success rate (ASR) alongside ASRD, its signed difference from a matched clean-trained control, to distinguish poisoning-induced effects from pre-existing target responses. In the rate study, LLM-CT achieves 95.9–100% ASR at 1% poisoning across BERT, GPT-2, and Llama-2-7B on SST-2, OLID, and AG News. Corresponding ASRD values of 92.9–98.6 percentage points demonstrate substantial poisoning-induced effects. Several baselines instead exhibit near-zero ASRD in some settings, indicating that their raw success largely reflects responses already present before poisoning. Fixed-trigger controls further demonstrate the value of low-confidence host selection at small budgets. Under seven individually applied BERT defenses, evaluated constructions retain 92.76–100% ASR at poisoning rates of 0.5–1%. Defense results show that partial poison removal and input rewriting can leave attack success high, motivating evaluation of post-defense behavior alongside detection rates and clean-data costs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.