mxSRBench: Towards Spatial Understanding and Diagrammatic Reasoning in mxGraph for Large Language Models
Abstract
While LLMs have advanced code synthesis and semantic reasoning, their ability to bridge the spatial–semantic gap in diagrams remains poorly understood. Diagrams demand satisfaction of semantic topology and metric geometry, yet existing benchmarks, focused on VQA or pixel generation, fail to isolate the resulting symbolic reasoning failures. We introduce mxSRBench, a benchmark of 3,990 instruction–XML pairs built on mxGraph, a high-fidelity symbolic language that explicitly encodes spatial logic. mxSRBench offers three tiers: T1 validates XML and structural renderability; T2 probes spatial reasoning; and T3 evaluates professional diagram synthesis across 11 types. For T2, we establish an Evaluation Protocol covering three core capabilities: (1) Understanding (interpreting spatial logic from XML), (2) Generation (grounding textual spatial constraints into valid coordinate XML), and (3) Multihop Reasoning (compositional multi-step spatial inference). Systematic benchmarking of LLMs/VLMs reveals a spatial–semantic gap: top models achieve up to 91.7% in understanding but degrade sharply on constrained generation and multihop reasoning, confirming they are sophisticated logicians yet poor drafters. mxSRBench provides a quantitative, diagnostic foundation for neuro-symbolic and tool-augmented diagram generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.