acceptodds
Under review as a conference paper at ICLR 2027

A GPU Environment for Chinese Standard Mahjong with a Judge-Replay Fidelity Contract and Reference RL Workloads

Abstract

Accelerator-native simulators have made reinforcement learning (RL) on board games cheap, but when a game's payoff is fixed by an external rulebook, a fast simulator is useful only if it is demonstrably faithful: a mis-scored hand still yields a plausible learning curve. We present a GPU-native JAX environment for Chinese Standard Mahjong (MCR), a four-player imperfect-information game in which a win is legal only if it scores at least eight fan (pattern points) in an 81-pattern catalogue, with a fidelity contract that replays 12,288 adjudicated competition games and requires exact agreement on every judge event, per-seat win-eligibility value and final score. The reference engine that the JAX environment mirrors passes all 12,288 games on every field and the JAX environment reproduces their events and final scores (not the win-eligibility values), each with a per-game record; we also report the corpus's rule coverage. We find that the simulator is not what limits RL training: it steps up to 6.9M transitions per second on one laptop GPU, while a training iteration spends 64% of its time in the policy network. Reference workloads show the benchmark is hard to learn. From scratch, reward shaping teaches agents to reach a ready hand, but legal win rates never exceed 2.8% and end at zero. Fine-tuning a strong imitation policy leaves its win rate unchanged and raises score per game by 0.33 points (95% CI [0.05, 0.60]), an exploratory result from one training run.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.