acceptodds
Under review as a conference paper at ICLR 2027

Learning from Population Marginals: Minimax Regret in Stackelberg Mean-Field Games

Abstract

In Stackelberg mean-field games, a leader interacting with a large population observes population states, while followers' actions are often hidden. We study reinforcement learning (RL) in such games with boundedly rational followers under state observation (SO), where the leader's population feedback consists only of the state distribution. Under SO, we establish the minimax regret over episodes, with other model parameters fixed, where is the effective dimension, is the follower feature dimension, is the number of leader actions, and are the numbers of follower states and actions. We develop an optimistic RL algorithm that attains this minimax rate by evaluating candidate follower responses using observed leader rewards and transitions, and a matching lower bound establishes the optimal dependence on . The exponent of increases with , quantifying how a larger effective dimension makes learning harder in the worst case. We then consider marginal observation (MO), which additionally reveals the action marginal distribution without state–action correspondence. When the SO rate is governed by the follower parameter dimension, we give a necessary and sufficient geometric condition for SO and MO to have the same minimax temporal order. We also establish matching separations for a family allowing arbitrarily large follower action spaces, showing when action marginal observation strictly improves the optimal regret rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.