KidVis: Benchmarking Fundamental Visual Primitives in Multimodal Large Language Models
Abstract
Despite the remarkable proficiency of Multimodal Large Language Models (MLLMs) in high-level reasoning, it remains questionable whether they possess human-like foundational visual abilities. To investigate this question, we present KidVis, a diagnostic benchmark grounded in cognitive developmental psychology rather than standard semantic recognition. KidVis deconstructs visual perception into six atomic primitives, including Visual Closure, Spatial Orientation, and Visual Tracking, and instantiates them through controlled, low-semantic, motor-free tasks. We evaluate 22 leading MLLMs against an age-specific human reference baseline. Our study uncovers a fundamental disconnect: while children achieve near-perfect accuracy of 95.31%, the strongest evaluated model, GPT-5, reaches 67.33%. We further observe a visual scaling paradox: within the evaluated model families, larger parameter counts do not consistently translate into stronger foundational visual perception, and may even coincide with semantic interference on low-level tasks. These findings suggest that current MLLMs still face systematic limitations in basic visual grounding, highlighting the need for evaluation and model design beyond semantic-heavy benchmarks and simple parameter scaling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.