MILD-MPC: Multi-Agent Imitation Learning via Distributed Model Predictive Control
Abstract
Controlling multi-agent systems in real time requires a scalable optimization method that satisfies constraints. However, existing distributed trajectory optimization methods typically assume fixed problem structures, and their solve times exceed the time budget as the number of agents grows. We propose Multi-agent Imitation Learning via Distributed Model Predictive Control (MILD-MPC), a framework for learning a permutation invariant predictive control policy that imitates a distributed MPC expert, generalizes to larger problems, and runs in a single forward pass at deployment. In addition, we develop Model Predictive Distributed Differential Dynamic Programming (MPD-DDP), a new distributed MPC method that handles time-varying problem structures, provides high-quality expert demonstrations, and serves as a backup policy when a safety verifier invalidates the learned policy's trajectory. We demonstrate that the deployed policy achieves substantial timing improvements over both a centralized MPC baseline and the MPD-DDP expert, generalizes to problems with more agents, and remains robust to agent failures during deployment, across 2D and 3D obstacle-field navigation tasks. In hardware experiments, a policy trained with 8 agents transfers to 15 robots without retraining, completing the task collision-free in real time.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.