acceptodds
Under review as a conference paper at ICLR 2027

Agentic Backdoors via Pretraining Poisoning

Abstract

Large language models are pretrained on vast, uncurated corpora scraped from the public web. This creates a security risk from attackers who plant poisoned content in the pretraining data. We show for the first time that such poisoning can backdoor an agent, causing it to execute attacker-chosen commands on the user's machine. To measure the threat end to end, we pretrain three Qwen3 models from scratch at sizes 0.6B, 1.7B, and 4B, and then post-train each into a CLI agent using a clean pipeline. Our main results show that poisoning only 0.1% of the pretraining corpus successfully implants the backdoor: The deployed agent executes a fixed malicious command when the trigger is present while behaving normally on clean inputs, and the attack success rate (ASR) strengthens with model scale, reaching roughly 92% success at 4B. We further introduce new poisoning designs that make both the poison document and the trigger harder to identify or filter out. We maintain 24% ASR even after reducing the poisoning rate to 0.0005%, only 2500 documents. We argue for treating the provenance of pretraining data as part of an agent's security model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.