acceptodds
Under review as a conference paper at ICLR 2027

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text

Abstract

Spatial intelligence enables agents to move beyond static semantic understanding and interact with the physical world. Many spatial tasks are grounded in continuous visual scenes, where locations, regions, and paths are more naturally expressed by pointing, marking, or drawing than by precise coordinates or discrete textual symbols. Yet existing spatial benchmarks usually require coordinates, options, or text, creating an answer-interface mismatch that makes it difficult to evaluate image-generation models under the same task semantics as text-output VLMs, despite their ability to externalize spatial judgments in pixel space. To address this problem, we propose ProVisE (Protocolized Visual Evaluation), a benchmark-agnostic framework that elicits protocol-constrained visual answers from image-generation models, parses them into structured predictions for task-specific scoring, and extends to previously unconfigured benchmarks through Agentic construction and validation of executable generation–parser protocols. We further introduce SpatialGen-Bench, a curated diagnostic benchmark comprising 470 samples across 14 spatial subtasks, four capability levels, and diverse answer forms. We evaluate representative text-output VLMs and image-generation models in a unified setting and validate Agentic protocol construction on six external spatial benchmarks. Experimental results show that image-generation models are competitive when spatial answers can be externalized in pixel space, while text-output VLMs retain a clear advantage in compositional spatial reasoning. These findings reveal complementary strengths of pixel-space expression and text-based reasoning, while ProVisE and SpatialGen-Bench establish a unified testbed for generative spatial evaluation and future studies of spatial cognition in image-generation models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.