A Solvable Model of Near-Constant-Count Data Poisoning
Abstract
Backdoor poisoning of language-model pretraining has been observed to succeed after a near-constant number of poisoned documents, largely independent of the amount of clean data, while packing more poison into each batch makes the attack less sample-efficient. Neither finding has a mechanistic explanation. We derive both in closed form. In a solvable model of online training, a trigger that never occurs in clean data () is learned from poison alone, so under SGD the threshold depends on the poison count and not on clean-data size; when the trigger also appears in clean data (), a mean-field fixed point recovers the fraction-governed regime of prior theory, and interpolates between the two. A pre-specified test on a transformer trained from scratch is consistent with this: the poison needed for a trigger that also appears in clean data tracks the predicted count of competing clean uses (123, 346 and 914 examples against a predicted 100, 300 and 900 as training length triples), while a trigger absent from clean data needs fewer than 40 at every length. Under Adam, which is how language models are trained, a second resource appears: because trigger-specific parameters receive gradient only on poisoned steps, each poisoned batch moves them by a fixed amount regardless of how many poison examples it contains, and batches closer together than the second-moment memory interfere. This zero-parameter prediction orders attack success across all 11 schedule, length and placement conditions we ran (Spearman , ), and a controlled optimizer intervention supports it: holding the total poison count fixed and splitting it across more batches raises attack success under AdamW (mean ASR at , ) but not under SGD (flat at ). Continuing pretraining of a pretrained model on natural text, with trigger frequencies measured from the corpus rather than imposed, reproduces the predicted ordering: the poison needed rises monotonically with corpus frequency and a never-occurring trigger is cheapest to implant by –. We also report the theory's limits: count invariance holds only within a window of training lengths, beyond which attack success decays in a way our account predicts in direction but underestimates in size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.