acceptodds
Under review as a conference paper at ICLR 2027

Sequential Data Poisoning in LLM Post-Training

Abstract

LLM post-training proceeds through multiple stages, e.g., supervised fine-tuning (SFT) followed by proximal policy optimization (PPO) or direct preference optimization (DPO), where each stage draws data from different, potentially untrusted sources. Existing literature assumes data poisoning attacks may occur at each training stage, but neglects the possibility of multi-stage attacks. To study the trustworthiness of the entire post-training pipeline, we propose the threat model of *sequential data poisoning*, where an adversary can poison both the SFT and preference datasets. Under this threat model, we identify the *single-attack illusion*: an attack at each stage, evaluated in isolation, appears to pose a weak threat. Yet when attacks compound across stages, the true vulnerability is revealed. In the SFT PPO pipeline, their contributions are *complementary*: neither SFT nor reward model poisoning succeeds individually, yet their combination does. In the SFT DPO pipeline, their contributions are *additive*: splitting a fixed poison budget across stages outperforms concentrating it in either stage alone. These findings show that security analyses of individual post-training stages systematically underestimate compound vulnerabilities that emerge only from their interaction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.