acceptodds
Under review as a conference paper at ICLR 2027

KidVis: Benchmarking Fundamental Visual Primitives in Multimodal Large Language Models

Abstract

Despite the remarkable proficiency of Multimodal Large Language Models (MLLMs) in high-level reasoning, it remains questionable whether they possess human-like foundational visual abilities. To investigate this question, we present KidVis, a diagnostic benchmark grounded in cognitive developmental psychology rather than standard semantic recognition. KidVis deconstructs visual perception into six atomic primitives, including Visual Closure, Spatial Orientation, and Visual Tracking, and instantiates them through controlled, low-semantic, motor-free tasks. We evaluate 22 leading MLLMs against an age-specific human reference baseline. Our study uncovers a fundamental disconnect: while children achieve near-perfect accuracy of 95.31%, the strongest evaluated model, GPT-5, reaches 67.33%. We further observe a visual scaling paradox: within the evaluated model families, larger parameter counts do not consistently translate into stronger foundational visual perception, and may even coincide with semantic interference on low-level tasks. These findings suggest that current MLLMs still face systematic limitations in basic visual grounding, highlighting the need for evaluation and model design beyond semantic-heavy benchmarks and simple parameter scaling.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.