MeanMAC: One-Step MeanFlow Policies with Team Q-Learning for Multi-Agent Coordination
Abstract
Offline multi-agent reinforcement learning (MARL) requires policies that can represent diverse cooperative behaviors from fixed datasets while remaining efficient enough for multi-agent decision making. Although diffusion- and flow-based policies provide substantially greater expressivity than conventional unimodal policies, existing generative MARL methods either rely on iterative sampling at execution time or adopt a two-stage pipeline that first learns a multi-step joint generator and subsequently distills it into one-step executable policies, thereby introducing additional inference or training overhead and separating generative modeling from the final policy. To address this challenge, we propose MeanMAC: One-Step MeanFlow Policies with Team Q-Learning for Multi-Agent Coordination, which directly learns a shared one-step MeanFlow policy for multi-agent action generation and couples policy improvement through a centralized team-value objective. The shared actor maps each agent's local information and noise directly to an action in a single forward pass, while the resulting joint action is optimized through team-level Q-learning, allowing generative distribution modeling and cooperative policy improvement within a unified training procedure. Consequently, MeanMAC avoids iterative denoising or ODE sampling and separate teacher–student distillation while retaining efficient factorized action generation and team-level coordination. Extensive experiments across three offline MARL benchmark suites, covering 10 environments and 33 task–dataset configurations on MA-MuJoCo, MPE, and SMACv1, demonstrate strong performance across both continuous and discrete control.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.