Maximin-Sufficient Learning under Strategic Feedback in Zero-Sum Games with an Informed Opponent
Abstract
A defender can test a policy in resettable diagnostic episodes before deployment, but an informed opponent may choose, among its best responses, those that hide the unknown game model. Identifying the model may then be impossible, and it is stronger than needed: the defender only needs one policy that is -maximin for every model the data cannot rule out. We ask when this weaker target can be certified. Because best responses can tie, each diagnostic query induces a set of transcript laws, from which the opponent selects adaptively. For a learner that sees only the recorded transcripts, we show that certification is impossible whenever a group of models that no single policy serves within tolerance shares a common transcript law at every accessible query. For finite models, queries and transcripts, with deployment and inference sharing one tolerance and small enough , this joint obstruction is the only one: a confidence sequence that tests empirical transcript frequencies against each model's set of laws returns an -maximin policy with probability at least exactly when no such group exists. An acquisition score from minimax-regret duality (RDI) targets the deployment decision and affects only speed. On matrix, patrolling, search and poker designs, identifying the model takes to times as many episodes as reaching the policy certificate on the same runs. On the patrolling and search designs, RDI query selection is about three times faster than uniform querying, mostly through an allocation that could be fixed before any data; the exact RDI score is at most times faster than a cheaper pairwise score there, and six times faster on an interdiction design built so that pairwise evidence vanishes. The certificate is conditional on the candidate library: in a misspecification study whose library misses the true game, the method certified every run, and % of the runs had a gap above the tolerance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.