Auditing Update-Order Effects in Multi-Agent Reinforcement Learning
Abstract
Update order can change a multi-agent reinforcement learning (MARL) algorithm even when its objective, data, and optimizer are fixed. We formulate an audit of the complete learning state, including optimizer memory, and characterize when an observable state remains sufficient under common future updates. The audit groups orders that are equivalent under commuting swaps and compares pairs of training trajectories after correcting for evaluation noise, retaining unresolved cases under explicit coverage assumptions. In three controlled PettingZoo/MPE environments, simplified independent-critic actor-critic updates have zero observed protocol range, whereas sequential shared-critic updates have ranges of 5.11-5.95 return points and two sample-mean ranking reversals. A critic-only probe can miss actor differences: actors read different intermediate states even when the final critic states agree. Freezing shared state restores invariance under disjoint actor updates. In HARL-based experiments at 10M steps per branch, reversing the native loop also reassigns random streams; agent-local streams isolate order and yield heterogeneous sensitivity estimates. Projecting finite update differences onto a surrogate utility gradient ranks held-out surrogate sensitivity (mean within-seed Spearman 0.78), but return-based scheduling and detector transfer remain unestablished. The resulting audit supports reproducible comparisons, with distribution-dependent small-seed inference; it does not establish a return-improving optimizer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.