Data Poisoning Attack against Batch Online Learners via Duplication and Deletion
Abstract
The recent trend in modern predictive models, such as LLMs, has prompted a change in the data source from curated datasets to a stream of real-time data. This change introduces a new vulnerability that an adversary could exploit to cause substantial damage while evading detection. However, the adversary is constrained in designing attacks based solely on observations available at decision time. Understanding the optimal attack strategy provides critical insights for designing defense mechanisms that secure predictive models and improve their resilience. In this paper, we study a white-box data poisoning attack against an online learner, where the adversary seeks to steer the learner's parameters toward a specific target. The adversary is restricted to deleting and duplicating existing observations within each incoming batch, without modifying their features and labels. We frame the attacker-defender interaction as a constrained Markov Decision Process (MDP) and propose a reinforcement learning (RL) agent that approximately derives the near-optimal attack. We provide theoretical analysis and numerical experiments to show that our algorithm achieves lower cumulative attack cost than a generic model-free RL baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.