acceptodds
Under review as a conference paper at ICLR 2027

Learning Coordinated Representations for Flow-Based Multi-Agent Imitation Learning

Abstract

Cooperative multi-agent systems require decentralized agents to coordinate seamlessly to accomplish shared tasks. Learning from multi-modal joint demonstrations that exhibit diverse strategies makes multi-agent imitation learning particularly challenging: independently executed policies may select individually plausible actions that accidentally mix distinct strategies, leading to inconsistent joint behaviors. We address this challenge by introducing the latent-conditioned factorization gap, which measures the divergence between a centralized joint policy and a product of latent-conditioned decentralized policies. We show that, with sufficiently expressive local policies, this gap decomposes into conditional total correlation among agents' actions and conditional mutual information terms capturing dependence on unobserved global information. To efficiently optimize this objective with flow-matching policies, we derive a practical, density-free surrogate based on joint-to-local vector field alignment. Extensive experiments on Multi-Agent MuJoCo and cooperative multi-quadruped benchmarks demonstrate that our approach significantly outperforms existing baselines in preserving coordinated multi-modal behavior.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.