COMRAD: A Benchmark for Embodied Multi-Agent Reinforcement Learning
Abstract
Benchmarks are central to the development and evaluation of multi-agent reinforcement learning (MARL) algorithms. As the cooperative MARL community has grown, two categories of evaluation environments have proven indispensable: low-dimensional feature-vector benchmarks that isolate algorithmic behavior in compact state spaces, and two-dimensional pixel-based benchmarks that rely on overhead visual observations. Embodied cooperation, however, requires high-dimensional 3D egocentric perception, partial observability arising from line-of-sight occlusion, and the ability to close the gap between recognizing a coordination pattern and executing it from a first-person view. While various 3D environments have been explored for vision-based RL, no existing platform simultaneously provides a standardized, purely cooperative multi-agent benchmark with 3D first-person observations and high-throughput simulation. To address this gap, we introduce , a operative ulti-Agent einforcement Lerning benchmark suite in oom, featuring a diverse set of challenging scenarios spanning role asymmetry, temporal synchronization, and spatial navigation. To introduce within-scenario variability, we develop , a procedural map generator that produces diverse layout configurations for every scenario. We integrate COMRAD with Sample Factory, a high-throughput asynchronous RL framework, and implement seven MARL baselines on top of it, reaching frames per second during training. Our experiments show that COMRAD poses significant challenges for current CTDE methods, and a capability-linked analysis of where each algorithm family fails, establishing visual cooperative MARL as an important open frontier.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.