acceptodds
Under review as a conference paper at ICLR 2027

Learning Search Guidance from Solver-Generated Ground Truth

Abstract

AlphaZero made strides in the field of game playing not just because it trained via self-play, but also because it learned *tabula rasa*, i.e., from a random initialization on self-generated data. That learning philosophy has not yet made its way into the field of game solving, which requires deriving the game-theoretic value of the initial position rather than merely selecting a good move. In this paper, we apply AlphaZero's learning philosophy in the context of game solving to introduce AlphaSolve, a game solver that learns search guidance from the game-theoretic values of intermediate nodes it proves while solving a game. We verify it on previously solved games: Connect Four and three Breakthrough variants. Programs for solving games are evaluated on their efficiency, which is determined by how well they search the game tree for a proof of the game-theoretic value. Prior art has relied on estimating proof cost (the number of positions required to solve a node), but we show that on Connect Four training on win/loss outcomes (as AlphaZero does) is competitive with training on the measured proof cost as the value target. AlphaSolve generates roughly a hundred times more ground truth labels than self-play game trajectories, and on Connect Four we find that a small share of trajectories mixed into the verdicts is better than either stream alone. This is significant because reinforcement learning algorithms are notorious for investing heavily in data generation, but AlphaSolve generates ground truth labels in abundance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.