Reinforcement Learning for Quantal Stackelberg Equilibrium in Mean-Field Games
Abstract
We study online reinforcement learning in episodic Stackelberg mean-field games, where a leader interacts with an infinite population of boundedly rational, myopic followers governed by a quantal-response model. The leader knows neither the followers' reward function nor their realized rewards and must infer their response model from behavioral observations. We analyze the complete observation (CO) setting, where the leader observes the exact joint state-action distribution of the population, and then extend the analysis to finite-sample observation (FSO), where only followers are observed per step and the true population distributions remain unobserved. We propose an algorithm combining kernelized optimistic least-squares value iteration with a confidence set based on maximum likelihood estimation for the followers' reward parameter. Under our stated assumptions, for a leader kernel of rank and a follower feature dimension , the algorithm achieves high-probability regret under FSO relative to the optimal CO policy, where is the number of episodes and is the horizon length. With fixed response and regularization parameters, this rate holds when , , and , where is the followers' state-action space. We further establish a CO leader-side lower bound , showing that the dependence is unavoidable. For the FSO commitment-rule class in this paper, we prove a fixed-dimensional lower bound ; hence is necessary for uniform expected regret within this class.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.