Certifying Tactical Mistakes in Self-Play Reinforcement Learning
Abstract
Self-play scores can conceal decisions that discard a forced win because sampled opponents fail to refute them. We present a certificate-based audit that turns local tactical proofs into executable opponent responses and quantitative guarantees on their value. The audit distinguishes mistakes on the realized deal from winning actions identifiable using the player's observations. Each denial witness supplies a payoff cap for a responding coalition, reducing certified response optimization to optimal stopping under ordinary self-play. A paired martingale-dual score bounds a fixed switching rule's regret within this response class, using exact conditional values or independent successor simulation. We give a sharp coefficient-one residual bound for predictor-greedy rules and an exact-boundary absorption result. In complete games of a sampled proximal policy optimization (PPO) policy, a simulated switching rule has mean certified score points per game, with stopping regret below points at 95% confidence. Across enumerable stochastic endgames, exact regret is and the paired dual bound is . Refutation replay denies first place on all audited throws. As a further use of the audit, reusing the certified action sets in training raises sampled retention from to on a fixed Tien Len challenge set over five independent runs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.