When Decisions Induce Policies: Performative Prediction with Latent Responses
Abstract
When a learned decision rule is deployed in an adaptive environment, future data are shaped by agents who respond to it. Performative prediction captures this feedback through induced distribution shift, but similar observable shifts can arise from different latent policy responses, which may require different interventions. We formulate this problem as a bilevel multi-agent reinforcement learning framework in which strategic users adapt to the incentives induced by a routing rule, while the institution updates the rule from aggregate feedback and behavior-derived evidence. To infer latent responses, we compare behavior windows against reference policies representing benign and targeted behavior. Theoretically, we connect user adaptation and platform learning through a governance-loss bound, and derive finite-time guarantees for predictive, stable user adaptation under explicit interaction and prediction-error conditions. Across competitive Riichi Mahjong and non-competitive platform environments, policy-evidence-based learning better distinguishes latent strategic responses than aggregate response signals, leading to more stable long-run regimes, fewer harmful users, and improved benign-user outcomes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.