acceptodds
Under review as a conference paper at ICLR 2027

Per-agent return distributions in value decomposition

Abstract

Distributional value decomposition for cooperative multi-agent reinforcement learning predicts a return distribution for each agent and combines the per-agent quantiles into a team distribution and risk-sensitive variants act on the per-agent distributions. We show that training on the team return does not identify them: at every state many decompositions give the same team quantiles and any decision rule that uses only team-level data violates a per-agent risk constraint with probability at least one half in some environment. Supervising each agent with its own return removes this ambiguity, but summing quantiles at matched levels, as DFAC-style mixers do, is comonotone recomposition, which for finite-variance returns overstates the variance of the team return unless agent returns are comonotone. We propose a critic that trains per-agent quantile heads on individual returns and predicts the team distribution with a separate head. In a heterogeneous search task and in Gigastep, team-only training left per-agent coverage far from nominal in most settings and dependent on the policy and on how credit was shared. The proposed critic lowered per-agent CRPS in every setting, kept team coverage at the team-only level where supervised matched-sum critics raised it and improved per-agent risk decisions in the search task.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.