Implicit Credit Learning for Sparse-Reward Multi-Agent Reinforcement Learning
Abstract
When a cooperative team receives feedback only at the end of an episode, a single outcome must guide many agents' earlier actions. Learning a dense team reward helps, but its decomposition does not automatically give each agent a useful policy update. We introduce IMAP, an online method that fits a shared Q/V reward model by comparing completed episodes according to their observed terminal outcomes. Local Q/V residuals provide agent-specific rewards; the same normalized weights combine them for a centralized critic and scale local advantages for decentralized PPO actors. We characterize when these local updates follow the learned team objective. With full returns and frozen reward heads, a sharp bound relates their gradient error to variation in future Q/V compatibility; an exact identity identifies additional continuation terms for general GAE. In experiments with terminal-only feedback, IMAP's local-actor configuration attains the highest reported mean among the compared methods on all nine SMACv2 scenarios and four MAMuJoCo tasks, and on two of three SMACLite scenarios. Controlled relay interventions confirm the predicted changes in actor credit while holding the learned team reward fixed. The gradient result concerns the fitted objective; benchmark returns measure performance on the original task outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.