acceptodds
Under review as a conference paper at ICLR 2027

RSD-Poker: Structure-Adaptive and Shift-Robust Risk–Utility Certification for Residual Policies in Imperfect-Information Games

Abstract

Imperfect-information reinforcement learning often requires adapting a strong base strategy to new opponents, environments, or deployment constraints, but directly optimizing residual policies can introduce unstable improvements and excessive risk under distribution shifts. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework for residual policy selection in imperfect-information games. RSD-Poker separates residual policy discovery from deployment-time certification by constructing adaptive evaluation structures, calibrating risk-utility tradeoffs, and selecting policies with statistical guarantees under diverse shift patterns. The framework supports stratified, partition-based, mixture-shift, learned-student, and opponent/solver evaluation endpoints through unified reporting and certification protocols. We evaluate RSD-Poker on imperfect-information poker environments under multiple distribution-shift settings. Compared with baseline residual adaptation strategies, RSD-Poker achieves improved utility while maintaining controlled exploitability risk. The Learned-Group Dual configuration obtains an average utility of 4.4936 with exploitability reduction of -0.0036, certification risk of 0.0078, and coverage of 0.7257. Under robust grouping, the method achieves competitive utility of 4.4827 while further reducing risk to 0.0059 and maintaining coverage of 0.7209. Results across stratified, partition, mixture-shift, learned-student, and opponent/solver endpoints demonstrate that adaptive certification enables reliable residual policy deployment beyond fixed-distribution evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.