Who Gets Blamed? Attribution-Evasive Adversarial Attacks in Cooperative Multi-Agent Reinforcement Learning
Abstract
Existing adversarial attacks in cooperative multi-agent reinforcement learning (c-MARL) mainly focus on degrading team performance, treating a lower return or win rate as the primary measure of attack success. However, after a team failure, post-hoc attribution analysis may assign responsibility among agents by estimating their relative contributions to the observed failure trajectory. An attack that successfully disrupts the team but makes the compromised agent highly attributable therefore leaves a strong diagnostic footprint. This motivates a distinct adversarial objective: a compromised agent should not only induce team failure, but also avoid becoming the most attributable cause of that failure. We study attribution-evasive adversarial attacks in cooperative multi-agent reinforcement learning and propose AE-APL (Attribution-Evasive Adversarial Policy Learning). The proposed framework augments a standard adversarial policy objective with an exposure-aware penalty, using a counterfactual Joint-Q attribution critic to estimate a local value-based surrogate of each agent’s harmful contribution. The resulting attribution scores are converted into a relative blame exposure signal, which penalizes the compromised agent when it becomes disproportionately responsible under attribution while preserving a controllable attack-stealth trade-off. Experiments on MaMuJoCo and SMAC show that AE-APL improves attribution stealth over strong attack baselines while maintaining competitive attack effectiveness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.