acceptodds
Under review as a conference paper at ICLR 2027

Toward Optimal Adversarial Attacks in Reinforcement Learning via Rate–Distortion Theory

Abstract

Reinforcement learning (RL) has been deployed in many security-related applications. To improve the robustness of RL systems, it is important to systematically study potential adversarial attacks. Most prior work has considered deterministic adversarial attack strategies targeting either the agent or the environment, which can often be mitigated or reversed by defensive RL agents. In this paper, we propose a provably "uncounterable" class of adversarial attacks on RL systems. Our attack leverages a rate-distortion information-theoretic approach to stochastically perturb the agents' observations of the underlying Markov Decision Process (MDP). These perturbations are designed to strongly or even entirely limit the agent's ability to obtain accurate information about the true environment. This proposed framework is modular and can be integrated with various observation attack strategies targeting different components of the MDP, including state, actions, transition functions, or reward signals. We derive an information-theoretic lower bound on the victim agent's reward regret and show the impact of rate-distortion attacks on state-of-the-art algorithms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.