acceptodds
Under review as a conference paper at ICLR 2027

One Model, Two Games: Certifying Nash Equilibria in Discounted Zero-Sum Low-Rank Markov Games

Abstract

An equilibrium of a learned game can admit profitable deviations in the true environment. We study how to certify approximate Nash equilibria in discounted zero-sum Markov games with unknown low-rank transitions. Our method solves two confidence games on one learned model. Under a policy-uniform two-sided value envelope and exact planning, pairing the lower game's maximizing policy with the upper game's minimizing policy minimizes the induced unilateral-deviation certificate. The reverse pairing controls its width and supplies sufficient exploration using the same solutions. A natural alternative combines a nominal equilibrium with a joint-action MDP maximizing bonus occupancy. We prove interaction upper bounds of the same order for both methods, while their fixed-envelope certificates can be strictly separated. Under finite-class realizability, known rewards in , reset access, and certified oracles, outer iterations suffice for an -Nash pair with high probability, where is the feature dimension and the joint-action count. Counting primitive transitions adds a factor . The guarantee allows controlled estimation and planning errors despite changing learned features. Experiments distinguish certificate tightness from policy quality and exploration cost: crossing gives tighter envelope certificates on shared data, whereas bonus and uniform exploration are faster on delayed-conflict games. The sufficient confidence parameters remain conservative at these budgets.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.