acceptodds
Under review as a conference paper at ICLR 2027

MatBak: Targeted-Behavior Adversarial Attacks on Multiple Reinforcement Agents via Preference-Based Intention Policies

Abstract

Adversarial attacks provide a crucial lens for evaluating the robustness of multi-agent reinforcement learning (MARL) systems, revealing worst-case failures and security risks in multi-agent systems (MAS) driven by deep reinforcement learning methods. However, existing attacks on MARL systems primarily aim to degrade cumulative returns, capturing generic performance failures while overlooking attacker-specified collective behaviors that are critical in safety-sensitive MAS. In this work, we study targeted-behavior adversarial attacks against MARL systems, where an attacker seeks to induce preferred team-level behaviors via bounded observation perturbations. Unlike standard reward-degradation attacks, targeted-behavior attacks require translating high-level attacker intentions into coordinated behavioral changes among interacting agents. We propose Multi-Agent Targeted-Behavior Adversarial Attack (MatBak), a preference-based framework for eliciting attacker-preferred collective behaviors in MARL systems. MatBak first learns an intention policy from attacker preferences. It employs a victim-observation encoder to evaluate attacker utility, then optimizes an observation-perturbation adversary via a dual-weighted bi-level mechanism measuring both agent-level engagement and state-level occupancy, enabling adaptive victim selection, targeted credit assignment, and fine-grained behavior manipulation. Across SMAC and MetaDrive tasks with MAT, MAPPO, and QMIX agents, MatBak achieves the highest attacker return in 26 of 27 settings; when used for adversarial training, it reduces average attacker return by 22.1% and improves average environment return by 11.4% over the strongest non-MatBak baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.