acceptodds
Under review as a conference paper at ICLR 2027

SpaEx: Target-Aligned Spatial Evidence for Multimodal Agents

Abstract

For spatial reasoning, multimodal agents built on vision language models (VLMs) must interpret measurements relative to the objects and reference frames specified by a goal. Aligned states still leave candidate comparisons to the model. We introduce SpaEx, a target-aligned interface coupling object-centered perception with explicit spatial execution and fusing candidate evidence with native VLM scores in the original answer space. With the VLM and perception tools fixed, the reader fits only four scalar parameters. Across SpatialMQA, SPAR-Bench single-view reasoning and CV-Bench 3D, SpaEx improves two backbones by 5.3–18.9 percentage points over matched task-adapted VLMs and attains the highest accuracy and semantic Macro-F1 among compared methods in every benchmark–backbone setting. Under shared observations, readers supplied with executed comparisons outperform those given aligned states even with longer training and compact inputs. Candidate-level fusion also recovers correct answers ranked first by neither branch. Although execution and fusion contribute differently across tasks, these results support explicit computation as an effective interface between spatial perception and native visual judgment.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.