Coordinated Value Gradients for Decentralized Test-Time Guidance in Offline MARL
Abstract
In offline multi-agent reinforcement learning (MARL), diffusion and flow policies capture multimodal joint behavior, but existing methods couple the generator to critic training, which limits their scalability and adaptability. Offline RL avoids this coupling through test-time critic guidance, but extending it to cooperative MARL is nontrivial: centralized critics require global inputs, whereas locally queryable critics need not improve team value. We propose Coordinated Q-Gradient Flow (CQF), to our knowledge the first test-time guidance framework for offline MARL, whose guidance is both locally queryable and team-aligned: monotonically factorizing a team-level critic's objective yields local utilities whose simultaneous guidance steps form a first-order ascent direction of the learned team value at every joint action. The flow policy is fit by supervised matching alone, and at deployment each agent steers it using only its observation and utility gradient. A one-hot relaxation extends the same pipeline to discrete control. Across MA-MuJoCo, MPE, SMACv1, and goal-conditioned MangoBench, CQF attains the best average scores among compared Gaussian, diffusion, and flow methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.