Practical Performative Policy Learning with Strategic Agents
Abstract
This paper studies performative policy learning, where agents adjust their features in response to a deployed policy to improve their potential outcomes, thereby inducing endogenous distribution shifts. While there is growing interest in strategic environments—including strategic classification hardt2016strategic and performative prediction perdomo2020performative—existing approaches often rely on restrictive parametric assumptions, such as micro-level utility models in strategic classification or macro-level distribution maps in performative prediction, which limit scalability. In this paper, we relax parametric assumptions on both micro-level agent behavior and macro-level data distribution. We identify a practical information structure in performative distribution shifts and introduce the evaluation vector—defined as the policy values at all feasible points a given agent can reach—as an effective mediator on the causal path from the deployed model to the shifted data. We propose a gradient-based policy optimization algorithm that uses a differentiable classifier as a surrogate for the high-dimensional distribution map. Our methodology exploits batch feedback, achieving superior sample efficiency over bandit feedback that summarizes responses into scalars. We provide theoretical convergence guarantees and demonstrate the method's efficacy in challenging high-dimensional settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.