MentalBlackboard: Evaluating Spatial Visualization via Mathematical Transformations
Abstract
Spatial visualization is the mental ability to imagine, transform, and manipulate the spatial characteristics of objects and actions. This intelligence is a part of human cognition in which actions and perception are connected on a mental level. To explore whether state-of-the-art Vision-Language Models (VLMs) exhibit this ability, we develop MentalBlackboard, an open-ended spatial visualization benchmark for Paper Folding and Hole Punching tests within two core tasks: prediction and planning. Our prediction experiments reveal that models struggle to apply symmetrical transformations, even when they correctly predict the sequence of unfolding steps. Additionally, rotations pose a significant challenge to models' physical situational awareness. The planning task reveals limitations in models' ability to analyze symmetrical relationships and implement the multi-stage symmetry process, with Claude 4.6 achieving the highest planning score at an accuracy of 14%. The top-performing model, o3, attains a peak performance of 71.6% on the generalization task, which does not require spatial visualization but transfers spatial data; however, it achieves only 25% accuracy on text-based prediction tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.