acceptodds
Under review as a conference paper at ICLR 2027

PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs

Abstract

When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small number of crafted instruction-response pairs that cause the model to embed attacker-specified entities, such as a year or location, in outputs for a target task while limiting their occurrence on non-target tasks. Prior work has demonstrated task-level poisoning attacks under individual experimental settings, but the lack of a common benchmark makes it difficult to systematically compare attacks and understand the factors that determine their success. We introduce PoisonForge, a benchmark that parameterizes this threat along four dimensions—bias type, poisoning mode, appearance count, and target length—and evaluate 12 open-weight models from five families, ranging from 2B to 32B parameters, under a small poisoning budget. With only 10 poisoned examples added to 1,000 benign examples, mean attack success rate (ASR) over 12 models and 16 poisoning configurations is \(38.8%\), with 11 of 12 models exceeding \(70%\) ASR in at least one configuration. We also show that unintended spillover to non-target tasks is low and performance on standard capability benchmarks is preserved. We analyze the factors that influence attack success and find that repeated appearances of poisoned content yield higher ASR than a single appearance, attack success depends on bias type, and longer target lengths generally reduce ASR.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.