PRISM: Peer-Regularized Induction via State Mining for Reward-Free Backdoor Attacks in Reinforcement Learning
Abstract
Backdoor attacks can manipulate reinforcement learning (RL) policies under hidden triggers while preserving benign behavior. Existing approaches mainly exploit victim-side information or manipulated training signals, overlooking the behavioral knowledge encoded in independently trained benign policies. We propose PRISM, a multi-policy-assisted backdoor framework that uses clean peer policies to guide selective poisoning. PRISM combines victim-side state salience with peer-based target-action compatibility to identify informative, low-conflict states, and performs reward-preserving target induction with peer regularization. Across visual control, safety control, and autonomous driving tasks, PRISM achieves strong attack success, preserves benign performance, and accelerates early-stage backdoor implantation. Our findings reveal a new attack surface in multi-policy RL: benign policy populations can unintentionally provide useful guidance for backdoor implantation.The anonymous implementation and experimental code are available at:https://anonymous.4open.science/r/ICLRproject-13C3
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.