acceptodds
Under review as a conference paper at ICLR 2027

GameSmith: Scaling Zero-Sum Self-Play with Synthesised Games

Abstract

Post-training language models on two-player zero-sum games provides a verifiable reward signal, and the resulting gains transfer to general reasoning. However, existing pipelines rely on a handful of hand-written games, with advantage estimators that are tied to those games. In this paper, we present GAMESMITH, a framework that scales this recipe on both fronts. To scale the supply of games, a language model synthesises games inside templates whose reward, observation, and hidden-information code is fixed. The generated game logic is verified by seed-paired replay and by an independent re-implementation of the same specification, and each verified game is then expanded into encoding, rule, and size variants. To scale post-training, GAMESMITH first distills a teacher’s self-play on the synthesised games into the model and then runs self-play reinforcement learning with CRAB, an advantage estimator that credits each decision against a shrinkage estimate of the expected return of its position and is defined on every variant of the corpus. Building on the three templates derived from the SPIRAL games, GAMESMITH produces 26 new games in 29 families, with 336 verified game variants. On Qwen3-4B-Base, distillation from the synthesised games raises the 13-benchmark reasoning average from 45.5 to 62.5, exceeding the same budget of the three SPIRAL games, and self-play with CRAB further improves it to 64.9. This ordering holds consistently across models, and transfer grows with the number of game families.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.