Toward Optimal Adversarial Attacks in Reinforcement Learning via Rate–Distortion Theory
Abstract
Reinforcement learning (RL) has been deployed in many security-related applications. To improve the robustness of RL systems, it is important to systematically study potential adversarial attacks. Most prior work has considered deterministic adversarial attack strategies targeting either the agent or the environment, which can often be mitigated or reversed by defensive RL agents. In this paper, we propose a provably "uncounterable" class of adversarial attacks on RL systems. Our attack leverages a rate-distortion information-theoretic approach to stochastically perturb the agents' observations of the underlying Markov Decision Process (MDP). These perturbations are designed to strongly or even entirely limit the agent's ability to obtain accurate information about the true environment. This proposed framework is modular and can be integrated with various observation attack strategies targeting different components of the MDP, including state, actions, transition functions, or reward signals. We derive an information-theoretic lower bound on the victim agent's reward regret and show the impact of rate-distortion attacks on state-of-the-art algorithms.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.