Optimal Joint-Action Screening for Cooperative Multi-Agent Reinforcement Learning
Abstract
Value factorization enables centralized training with decentralized execution in cooperative multi-agent reinforcement learning, but the widely adopted monotonic mixing structure leads to the suppression of individually optimal actions when they are paired with inferior teammate actions. Existing weighted methods mitigate this interference by identifying and emphasizing potentially optimal joint actions. However, inaccurate optimality criteria may exclude truly optimal samples from the protected set while including non-optimal ones, undermining the intended effect of sample weighting. Instead of identifying optimal joint actions directly, we seek a screening threshold that retains all optimal joint actions while safely identifying non-optimal samples for downweighting. We show that the maximum value from the least-squares projection onto the QMIX function class lower-bounds the optimal joint-action value and therefore provides such a threshold. Based on this, we propose Optimal Joint-Action Screening (OJS), which uses an independently trained reference QMIX to construct the candidate set. Candidate samples receive unit weight, while the remaining samples are adaptively soft-weighted according to their potential interference with optimal-action learning. Experiments across diverse cooperative benchmarks show that OJS achieves competitive or superior performance over representative value-factorization baselines.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.