acceptodds
Under review as a conference paper at ICLR 2027

Local and Global Stability in Performative Reinforcement Learning

Abstract

In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a *performatively stable* policy that is optimal in the environment it induces. Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map, which are hard to verify and fail in settings such as multi-agent best-response dynamics. We instead study stability for *mixtures* of policies, and show that the resulting picture is fundamentally different from performative prediction, where randomization removes the need for any sensitivity assumption. We distinguish *local* mixed stability, an occupancy-weighted first-order relaxation, from *global* mixed stability, which certifies against arbitrary deviating policies. The two notions genuinely differ: we exhibit an instance where local stability is achieved exactly but every mixture has global stability gap bounded away from zero. We show that a weighted per-state Hedge dynamic drives the local stability gap to zero at an rate for an *arbitrary*, possibly discontinuous, environment map. For global stability we introduce a weaker bounded transition range () assumption, under which unweighted per-state Hedge converges up to a floor of . This floor is unavoidable as we prove a matching-in- lower bound under trajectory feedback. Finally, we extend both notions to -player Markov games, obtaining local stability with no assumption on the environment map or game structure, and global stability for Markov potential games.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.