acceptodds
Under review as a conference paper at ICLR 2027

Certifying Exploitability at Learned Equilibria in Multi-player Games

Abstract

We develop an amortized exploitability certificate for fixed points of learned response maps in games with compact convex product action sets and costs convex in each player's own action. We certify the exploitability of arbitrary strategy profiles with playerwise envelopes without requiring contraction, monotonicity, or uniqueness of equilibrium. Its two game-dependent quantities are each player's own-gradient norm and the consistency error between the learned response operator and the pseudo-gradient. Both are enveloped once by cost-oracle probes on a finite -covering of the certified region . Subsequent profiles require only a forward pass. Two results delimit the scheme. Amortization is necessary: at a profile where every player's projection is inactive, a single probe cannot improve the model-free bound. Amortization is also costly: coverings at fixed accuracy grow exponentially with the dimension of . Across the two product-set games and seven response families, the certificate is valid in all theorem-supported cells; we report coupled-constraint cells separately as an out-of-scope stress diagnostic. Small-gain contraction certification is neither necessary nor sufficient for low exploitability in the models tested. Finally, on a nonconvex game the natural residual vanishes analytically at a profile of exploitability , so that playerwise convexity cannot be dropped.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.