SRC: SINKHORN RESPONSIBILITY CREDIT FOR SPATIAL MULTI-AGENT REINFORCEMENT LEARNING
Abstract
Credit assignment in cooperative multi-agent reinforcement learning (MARL) is commonly formulated as attributing a shared team return to individual agents through value estimates, counterfactual baselines, or decomposed utilities. In spatial and object-centric tasks, however, coordination also depends on which agents should cover which landmarks, goals, resources, or targets. Shared rewards and global shaping signals can provide useful team-level feedback while leaving this assignment pattern unresolved. We propose Sinkhorn Responsibility Credit (SRC), a transport-based credit module that models this pattern as an entropic optimal-transport plan between agents and task objects. SRC retains the plan as a soft responsibility matrix and converts it into dense agent-specific credit. This differs from global optimal-transport shaping, which uses the same geometric objective but broadcasts a single team-level scalar. Under a fixed policy backbone, marginal SRC improves MPE simple spread returns relative to shared rewards by 5.07 to 20.03 points and relative to global optimal-transport shaping by 2.71 to 16.50 points, with larger gains as the number of agents grows. This isolates the central mechanism: the same spatial objective is more informative when its transport assignment is returned to individual agents than when it is collapsed into a global potential. On VMAS navigation, static and temporal SRC scale more reliably than leave-one-out marginal responsibility: across N=4,6,8 and 20 matched seeds, SRC-S wins 60/60 and SRC-T wins 59/60 seed-wise comparisons against shared rewards with paired p<7×10^-8 at every team size. Difference rewards remain the strongest MPE baseline because they evaluate task-utility counterfactuals directly; SRC instead provides an assignment signal from observed agent-object geometry. Auxiliary learner checks, hybrid experiments, BenchMARL/TorchRL MAPPO runs, and stress tests further identify the regime in which SRC is most effective: tasks where observable objects and meaningful costs expose decision-relevant assignment structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.