PGV: A Two-Check Verification Harness for Self-Play Promotion Gates, Validated on Four Real Incidents in Deployed Systems
Abstract
Competitive self-play reinforcement learning relies on a promotion gate to decide which candidate policy is actually better. When the gate's own measurement is wrong in a way that is not obviously noise, the loop can silently stall, silently promote a policy that has not improved, or silently reuse data it presents as independent. We introduce PGV, a lightweight verification harness of two offline checks for self-play promotion gates, each a one-line invariant over bookkeeping the gate already keeps: the Disjoint-Reseed Check (a candidate's certifying data, and any data offered as an independent replication, must occupy regions of the seed space disjoint from the selection sweep and from each other) and the Train/Eval Symmetry Check (every structural degree of freedom the evaluation condition exposes - seat, side, deck - must be sampled the way training samples it). Neither check was run in either of two independently built, deployed self-play systems we studied. We validate PGV on four real incidents rather than a synthetic benchmark. The Disjoint-Reseed Check flags two real certification failures - a candidate certified as best of roughly twenty with a +0.0054 placement-point edge that falls to -0.0005 (interval spanning zero) under selection-disjoint fresh evaluation, consistent with a winner's curse, and a seed-stepping bug that made a nominal "independent replication" silently reuse 58.33% of the original run's blocks - and, applied to a different, later improvement claim in the same system, passes: five mutually disjoint evaluations all clear the calibrated tie point, so the check does not drive every claim to null (of two independently trained replicas, one clears on its own; their pooled estimate clears). The Train/Eval Symmetry Check would have flagged the imbalanced gate condition behind a 124-generation freeze of an independently built card-game ladder, diagnosed from a seat-conditional win-rate split (0.14 vs. 0.66); the mahjong gate avoids the identical mismatch by explicit design, and a pre-registered synthetic reproduction shows the mechanism's direction only at larger biases while falsifying its own pre-registered magnitude threshold, reported as falsified. Every number traces to a committed file or a pinned upstream commit, and an evidence plan frozen before the analyses fixed how each claim is narrowed when its own check comes back negative, which happened twice and is reported both times.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.