acceptodds
Under review as a conference paper at ICLR 2027

Pretraining Data Can Be Poisoned via Computational Propaganda

Abstract

We demonstrate the feasibility of poisoning attacks on the massive and heterogeneous open web via the existing infrastructure of webpages. Specifically, we show that third-party content injection through public discussion interfaces can scalably poison scraped web data, the dominant data source of modern pretraining. To do this, we introduce HalfLife, a measurement framework for estimating pretraining data vulnerability through analysis of injection opportunities and poison survival through data pipelines. Applying HalfLife to Common Crawl webpages, we estimate that up to 3.5% of final web training documents can be affected by this poison. At web scale, this can affect more documents than all of Wikipedia as a contributing data source. Finally, we pretrain a ladder of 65M to 1.3B parameter models on web pages with small rates of comment poisoning, and find consistent shifts in base model preferences across scale that can persist after instruction-tuning. The feasibility of poisoning beyond the limited data sources explored by prior work emphasizes the threat that pretraining data poisoning can pose, especially since attacks can be difficult to detect and potentially irrecoverable.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.