acceptodds
Under review as a conference paper at ICLR 2027

Calibrating Policy Covers for Nash Equilibrium Learning

Abstract

A policy cover can reach every important state while its uniform mixture provides poor data for equilibrium learning. We study how to calibrate a finite cover in zero-sum Block games without identifying the latent-state decoder. Under common emissions and a finite evaluable decoder class containing the true decoder, constraints on all candidate-decoder events are equivalent at the population level to latent coverage constraints. Incorrect candidates therefore do not require a more conservative population mixture. Every library admits a coverage factor R\leq\min \ {M,S\}, where is its size and the latent state count. We estimate the event constraints from shared rollouts and implement them by linear programming. Under finite-class fitting conditions, the resulting Nash guarantee replaces by in the sufficient fitting budget, while charging cover construction and calibration. A hidden-coordinate class gives a polynomial-time implementation. In a three-stage family with an unknown entrance, episodes match an adaptive lower bound up to logarithms, whereas uniform sampling of the same library requires . Fixed-budget experiments distinguish distribution quality from the costs and best-response errors that determine downstream performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.