acceptodds
Under review as a conference paper at ICLR 2027

Noisy Replicator-Mutator Dynamics in Multi-Agent Learning

Abstract

Evolutionary game theory predicts where a population of learning agents ends up, but the classical replicator equation is derived for a payoff matrix that does not change. Two recent lines of work relax it in different ways. Wang et al. (2026) show that global payoff noise adds a multiplicative term , with the intensity of the payoff fluctuation, to the replicator equation, stabilising coexistence and generating stable limit cycles; Bauer et al. (2026) show that an additive mutation term turns it into a replicator-mutator system whose interior equilibria are last-iterate attracting, which yields convergence guarantees for multi-agent reinforcement learning. The two perturbations have been analysed on their own; their interaction has not been studied. We derive the system that carries both, which we call Noisy Replicator-Mutator Dynamics (NRMD), from a stochastic policy-gradient update; the two terms act on different quantities, so they can be set independently. Our contributions are threefold: (i) we give a complete two-strategy theory of NRMD, consistent with its -strategy form at ; (ii) we validate the effect of noise in multi-agent experiments; and (iii) we propose a novel algorithm, Population-Replicator Policy Gradient (PRPG), that exploits NRMD. Code and data will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.