Stationary Alone, Drifting Together: Adapting to Interaction-Mediated Nonstationarity in Multi-Agent Reinforcement Learning
Abstract
External conditions can change how agents interact through shared infrastructure without changing their dynamics in isolation. This drift persists with frozen partners and vanishes for a lone agent. We formalise this interaction-mediated nonstationarity as a family of partially observable Markov games indexed by severity and introduce three public benchmarks for routing, voltage control and cooperative lifting, preserving their original rewards. We propose PACT (Peer-Action Coupling Tracking), a wrapper for unmodified multi-agent learners. Each agent predicts disturbances from peers' broadcast actions and the infrastructure layout using an online model whose size is independent of fleet size. A trust level controls compensation or steering, with zero trust recovering the host exactly. In a linearised feeder, predictive compensation with exact estimates raises the instability threshold without bound as trust approaches its maximum. In a routing model, excessive trust causes oscillations that make the route predicted to be faster become slower. Experiments against representative baselines demonstrate consistent improvements over the host learners across all benchmarks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.