Comparing LLM Performance on Chess Variants
Abstract
Recently, frontier LLMs have achieved rapid capability gains on reasoning-intensive benchmarks including math and coding. However, recent work has indicated that many of those capabilities may come from memorizing expanded training data rather than general gains in reasoning ability. To assess this possibility, we use chess as a testbed and test the performance of LLMs on a variety of non-standard chess variants, which are less represented in the training data than standard chess, against the chess variant engine Fairy Stockfish. We find that standard chess performance is strong in the opening moves and degrades in the midgame, while performance on chess variants remains weak throughout the game, providing support for the claim that memorization/overfitting to the training distribution accounts for a sizable portion of the high performance on reasoning benchmarks, which is further supported by our observation that LLMs perform better when given input chess moves in algebraic notation as compared to FEN. We also test a variety of other settings and find that enabling long chain-of-thought offers only marginal performance improvement in the midgame, and play against a “human-like” opponent or a non-human-like opponent of the same skill does not significantly affect performance. Overall, LLM performance is strongest in the settings most commonly represented in training data, and weaker on less represented settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.