What Counts as Evidence for a World Model? Reassessing Othello-GPT
Abstract
Othello-GPT, a transformer trained to predict legal moves in Othello, has become the standard example of a neural network that uses a world model. The evidence falls into three categories: near-perfect prediction of legal moves, near-perfect linear decoding of the board state, and highly successful causal interventions. We argue that this evidence is ambiguous once we adopt the right baseline and the right metrics. Our baseline, which we call Othello-MLP, uses heuristics rather than a world model, and scores nearly as well on all three criteria. We then introduce an additional piece of evidence—transfer learning speeds on geometrically coherent versus incoherent extensions of Othello—which further suggests that Othello-GPT is using a collection of heuristics. We conclude that the total evidence is at best ambiguous, and if anything favors a heuristic interpretation. We draw four general lessons. Two concern how to test for world models: (1) when attributing a world model, the right baseline is a network that uses heuristics rather than a network with randomized weights; and (2) the right metrics take into account patterns of error and the full output distribution. The other two lessons are about how to weigh the different kinds of evidence: (3) causal interventions can provide illusory evidence; and (4) transfer learning can be more revealing than task performance, decoding accuracy, and causal interventions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.