When AlphaGo Thinks: Mastering Go with an LLM-based Reasoning System
Abstract
Benefiting from the rapid development of reinforcement learning with verifiable rewards (RLVR), reasoning LLMs have shown remarkable capabilities on a wide range of reasoning tasks, especially in mathematics and code. However, when the focus shifts to reasoning tasks with scarce textual corpora, such as Go, their performance remains far from satisfactory. Although recent studies have trained LLMs that approach the level of professional Go players, these models still struggle to produce precise professional Go analysis. In this work, we investigate the possibility of building a Go-domain reasoning system based on LLMs. We first examine the capability boundaries of LLMs of different sizes when they serve as policy models and reward models for Go tasks. Based on these findings, we build an agent system consisting of a sampler and a reward model. We further propose Term-DPO, an algorithm that substantially improves the model's accuracy in expressing Go-related professional expressions. Through extensive experiments, we build the first LLM-based reasoning system that reaches the capability level of specialized Go models and surpasses top human players. At the same time, the system provides interpretable analyses with accurate Go terminology.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.