Go-MLLM-Bench: Perception, Decisions, and State Updates in 1919 Go
Abstract
Multimodal models are increasingly expected to act on what they see, but reading a scene, choosing an action from it and updating it after the action are different skills that one success rate mixes together. We introduce Go-MLLM-Bench, which asks about the same 1919 Go positions through an image, an exact board matrix or the move history, and scores board reading, move prediction and board updates separately. We will release nested training cohorts from disjoint games with board, engine and mixed targets. The main finding is simple. Models can learn to read the board almost perfectly, but this skill does not carry over to choosing moves or to updating the board after a move. A fine-tuned model that reads nearly every tracking image exactly never produces an exact update, even from the true previous board with a decoder that forces valid output. Handing the update to a rule engine, starting from the model's own reading, recovers every state. The benchmark also shows that locating a salient mark is not reading the board, that cell accuracy rewards copying a stale board, and that legal-move decoding lengthens games without finishing any. Benchmarks should measure what a model does with a state, not only whether it can report it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.