acceptodds
Under review as a conference paper at ICLR 2027

Beyond a Single Drawing: MultiFloorQA for Cross-View Spatial Reasoning in Architectural Drawings

Abstract

Spatial reasoning in multimodal large language models (MLLMs) remains insufficiently understood, particularly when spatial evidence must be integrated across multiple views. MultiFloorQA is introduced as a benchmark of 608 question–answer pairs over multiple architectural drawings, spanning six spatial task categories and Single-floor, Cross-floor, and Cross-projection evidence. Drawing generation, question construction, and reference-answer computation are grounded in a shared building representation, enabling precise and reproducible evaluation. Evaluation of representative MLLMs reveals substantial variation across models and spatial operations, while precise geometric computation and larger-scale spatial composition remain challenging even for strong systems. Controlled input-condition experiments further indicate that recovering and organizing geometry from the drawings themselves is a major source of difficulty. These findings motivate further advances in visual grounding, geometric computation, and cross-view spatial integration.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.