Pessimistic Follower Responses for Robust Stackelberg Reinforcement Learning
Abstract
Stackelberg and bilevel reinforcement learning optimize a leader while accounting for the adaptation of a follower, but typically assume that a particular optimal or entropy-regularized follower response will be realized. This can be fragile when approximately optimal follower behaviors have substantially different consequences for the leader. We introduce a locally pessimistic response model for two-player Stackelberg Markov games that exponentially tilts the follower's entropy-regularized best response toward actions that are less favorable to the leader, without modifying the follower's learning objective. We derive an off-policy hypergradient that differentiates through both follower adaptation and the pessimistic response transformation. We show that the proposed response is exactly worst-case within an induced local KL neighborhood, remains -optimal for the follower, and that optimizing the leader against this pessimistic response sacrifices at most nominal leader performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.