Identification and Off-Policy Evaluation in Markov Games with Adaptive Opponents
Abstract
Deploying a policy in a multi-agent system can change how the other agents behave. Off-policy evaluation should therefore measure the policy against the opponent response induced by deployment, not against an average opponent in the log. We show that this induced-response value is unidentified without observable response structure and cannot be estimated when the target lacks support on relevant state-joint-action triples or state-response pairs. We introduce a finite response-state model in which the response at stage depends only on earlier response states. The induced path is therefore obtained by forward recursion, not by solving a simultaneous fixed point. Given analyst-specified response maps , our polynomial-time method reconstructs response labels for the logging policies, estimates an opponent-response kernel, recovers the response to each undeployed target in one forward pass, and applies pessimistic off-policy evaluation on the target-reachable support. When overlap is incomplete, we give outer confidence intervals under partial identification for completions that share a response path. We also bound the value error caused by response-model misspecification. Together, these results distinguish structural identification from response recovery and statistical uncertainty.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.