SpatialCanvas: A Visual Canvas Harness for Explicit Spatial Reasoning
Abstract
Vision-language models can recognize local spatial relations yet still fail on compositional spatial questions that require reasoning across viewpoints or reference frames. We argue that a key failure mode is not missing visual evidence itself, but losing the association between an observation and the entity, source view, and reference frame under which it is valid. We introduce SpatialCanvas, a training-free inference framework that makes these bindings explicit for a frozen VLM. Given a spatial question, SpatialCanvas compiles the required entities, reference frames, spatial operations, and evidence dependencies into a typed program. The VLM supplies source-bound observations, while SpatialCanvas instantiates the corresponding frames, stores observations and derived relations in an image-bound state, and executes only frame-compatible spatial transformations. A structural validator checks whether the evidence required by the program is supported before producing the final answer. With frozen Qwen3-VL-4B, SpatialCanvas improves accuracy over direct inference from 28.1 to 30.5 on MMSI-Bench, from 34.9 to 44.4 on MindCube-tiny, and from 45.7 to 48.1 on OmniSpatial. These results suggest that explicitly preserving reference-frame and source-view bindings can improve compositional spatial reasoning without updating model parameters. Our code will be made publicly available soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.