acceptodds
Under review as a conference paper at ICLR 2027

Self-Play Evaluation Under Population Exposure Mismatch: Unidentified Verdicts on Defensive Skill in Competitive Mahjong

Abstract

Game-playing agents are usually selected in-house, against copies of themselves or ancestors, before they meet their deployment population. We study one way this selection can mislead: a skill may pay off only in states that opponents create, and the in-house population may create those states at a different rate. We measure this population-exposure mismatch for a Chinese Standard Mahjong agent, using deal-in incidence as a policy-dependent proxy for threat-state exposure, and ask what an in-house verdict on a defensive intervention identifies. With our simulator's scoring defect corrected, our strongest in-house population exposes the deployed policy at times the live-field rate ( CI –) and at a competition final's rate. Under stated assumptions for transporting a verdict, any exposure gap leaves the deployed sign of a fold policy with a precise negative in-house verdict unidentified: against the live field the sign flips if an exposed game gains over placement points; against the final the verdict carries over. Gate verdicts predate the correction. A live A/B resolves neither the deal-in effect nor placement, which would take thousands of games per arm. In that final, the top two agents' statistical tie splits into opposing score channels; finalists face threat states equally often but differ in how often a threatened discard deals in. We distil an evaluation recipe that separates exposure-validity checks from general hygiene. All evidence comes from one game and one agent lineage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.