Causal Shapley Interaction Consistency for Credit Assignment in Cooperative Multi-Agent Reinforcement Learning
Abstract
Shapley-based multi-agent reinforcement learning (MARL) methods distribute a shared reward according to agents' marginal contributions to provide a credit assignment as a prominent topic in artificial intelligence research. However, existing methods often overlook the relationship of causal semantics between agents' actions, leading to inaccurate credit assignment. To address this issue, we propose a Causal Shapley Interventional Consistency Regularization framework, namely THEMIS, which separates causal identification, critic calibration, and policy optimization through three components. The causal-identification component applies a one-time direct intervention to recognize causal relationship between agents' actions. The critic-calibration component applies an interaction consistency to align selected critic-induced interactions with these causal targets. The actor component applies a detached conventional advantage to prevent teacher gradients from influencing the policy update. Across different agents MPE tasks, THEMIS exceeds the normalized team-reward baselines on both learning metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.