acceptodds
Under review as a conference paper at ICLR 2027

MATCHA: Bridging the Gap between Sample and Wall-Clock Efficiency in Model-Based Multi-Agent Reinforcement Learning

Abstract

Model-based Multi-Agent Reinforcement Learning (MBMARL) has demonstrated superior sample efficiency compared to model-free baselines. However, its practical impact is often limited by high wall-clock training time. To bridge this gap, we introduce **MATCHA**, a system-algorithm co-design that makes search-based MBMARL efficient on modern accelerators. At system level, MATCHA implements an end-to-end, device-centric framework that keeps data collection, planning, and learning on the accelerator, massively reducing costly CPU–GPU communication. At the algorithmic level, we propose Path-Guided Frontier Search (PGFS), which selects nodes directly from the frontier of the tree using a global score based on cumulative path statistics. This reduces selection and expansion to a single vectorized tensor operation, avoiding depth-dependent sequential tree traversal on the GPU. Across SMAX, Overcooked, and MPE, MATCHA matches or exceeds the sample efficiency of strong MBMARL baselines while substantially reducing training time, achieving up to end-to-end wall-clock speedup with minimal GPU resources.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.