PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs
Abstract
When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small number of crafted instruction-response pairs that cause the model to embed attacker-specified entities, such as a year or location, in outputs for a target task while limiting their occurrence on non-target tasks. Prior work has demonstrated task-level poisoning attacks under individual experimental settings, but the lack of a common benchmark makes it difficult to systematically compare attacks and understand the factors that determine their success. We introduce PoisonForge, a benchmark that parameterizes this threat along four dimensions—bias type, poisoning mode, appearance count, and target length—and evaluate 12 open-weight models from five families, ranging from 2B to 32B parameters, under a small poisoning budget. With only 10 poisoned examples added to 1,000 benign examples, mean attack success rate (ASR) over 12 models and 16 poisoning configurations is \(38.8%\), with 11 of 12 models exceeding \(70%\) ASR in at least one configuration. We also show that unintended spillover to non-target tasks is low and performance on standard capability benchmarks is preserved. We analyze the factors that influence attack success and find that repeated appearances of poisoned content yield higher ASR than a single appearance, attack success depends on bias type, and longer target lengths generally reduce ASR.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.