acceptodds
Under review as a conference paper at ICLR 2027

A Semi-Decentralized Approach to Scalable Multiagent Control

Abstract

Scalable multiagent control under stochastic communication remains an open challenge. Existing planning methods such as RS-SDA scale poorly, sampling-based approaches such as Dec-MCTS assume fixed-period communication, and MARL methods typically impose either strict execution-time decentralization (QMIX, MAPPO) or full centralization (MAZero), respectively limiting coordination opportunities or ignoring the stochastic communication constraints that arise at decision time. We consider the Semi-Decentralized POMDP (SDecPOMDP), which unifies the Dec-POMDP and MPOMDP through a distribution over communication-sojourn times. We introduce SDecMCTS, a multiagent MCTS algorithm that achieves solution quality comparable to RS-SDA while requiring less computation across benchmarks. Building on the SDecPOMDP's semi-Markov communication dynamics, SDecMCTS uses an options framework in which decentralized solvers execute between communication events, enabling planning and execution that faithfully reflects semi-decentralization. We extend SDecMCTS to Semi-Decentralized Zero (SDecZero), a scalable AlphaZero-style algorithm that learns semi-decentralized value estimates used by decentralized solvers. SDecZero additionally learns the stochastic communication dynamics through a dedicated network communication head. SDecZero substantially outperforms model-based and model-free methods across large multi-drone search scenarios constructed from real cave systems associated with the DARPA Subterranean Challenge.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.