Adversarially Robust Multi-Agent RL via Test-Time Imagination
Abstract
Recent advances in model-free multi-agent reinforcement learning (MARL) have improved policy robustness against adversarial observation perturbations at test time. However, training such robust policies requires extensive interaction with the environment. Multi-agent world models could reduce the interaction cost by enabling robust policy learning from imagined experience. Multi-agent world models, however, typically condition their predictions on global information unavailable to individual agents at execution time, making them unsuitable for decentralized defense. Moreover, policies trained from imagined experience are not inherently robust to adversarial observation perturbations. We propose MATRIX, a world-model-based MARL framework that unifies sample-efficient policy learning with training- and test-time mechanisms for robustness to test-time adversarial observations. MATRIX applies classifier-free guidance to a diffusion-based world model that predicts each agent’s next observation using either local information alone or using local and global information jointly. This design enables MATRIX to use a single world model for three complementary functions: training policies on model-generated trajectories, robustifying policies using model-generated trajectories, and predicting observations from local information for decentralized test-time defense. Building on locally conditioned observation prediction, we propose a low-overhead execution protocol that uses the trained world model to mitigate the effects of adversarial observation perturbations at test time, without requiring a separate defense model. Evaluations in environments with discrete and continuous action spaces demonstrate that MATRIX outperforms state-of-the-art baselines in both sample efficiency and robustness to test-time observation attacks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.