acceptodds
Under review as a conference paper at ICLR 2027

Coordinated Value Gradients for Decentralized Test-Time Guidance in Offline MARL

Abstract

In offline multi-agent reinforcement learning (MARL), diffusion and flow policies capture multimodal joint behavior, but existing methods couple the generator to critic training, which limits their scalability and adaptability. Offline RL avoids this coupling through test-time critic guidance, but extending it to cooperative MARL is nontrivial: centralized critics require global inputs, whereas locally queryable critics need not improve team value. We propose Coordinated Q-Gradient Flow (CQF), to our knowledge the first test-time guidance framework for offline MARL, whose guidance is both locally queryable and team-aligned: monotonically factorizing a team-level critic's objective yields local utilities whose simultaneous guidance steps form a first-order ascent direction of the learned team value at every joint action. The flow policy is fit by supervised matching alone, and at deployment each agent steers it using only its observation and utility gradient. A one-hot relaxation extends the same pipeline to discrete control. Across MA-MuJoCo, MPE, SMACv1, and goal-conditioned MangoBench, CQF attains the best average scores among compared Gaussian, diffusion, and flow methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.