What Teammates Learn Next Changes Which Updates Are Worth Making
Abstract
Off-policy cooperative multi-agent reinforcement learning (MARL) reuses past interactions through experience replay, but limited update budgets permit learning from only a fraction of candidates, while each update's team utility depends on teammates' subsequent learning and responses. However, existing replay and data-attribution criteria largely use current policies and error signals, without accounting for subsequent teammate learning. The most valuable update with teammates held fixed need not remain optimal under co-adaptation, a phenomenon we call Co-Adaptive Replay-Value Misalignment (CoRVM). To address CoRVM, we propose ExSpectra, which estimates co-adaptive update utility before replay allocation. Co-Adaptive Utility Decomposition (CAUD) exactly decomposes co-adaptation-induced utility changes under fixed-data matched interventions into baseline teammate-learning and update-induced teammate-response effects. Based on CAUD, we construct a first-order local influence estimator that augments frozen-teammate utility with co-adaptive corrections using cross-agent evaluation curvature and update sensitivity. Shared Hessian–vector products and cross-agent curvature estimates amortize second-order computation across candidates, while sparse paired probes calibrate residual error. Held-out matched audits across five benchmark families confirm that CoRVM changes finite-budget replay decisions. Extensive experiments show that ExSpectra outperforms strong replay-based and resource-matched baselines on most tasks. Code is anonymously available at https://anonymous.4open.science/r/ExSpectra-0DC0/
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.