acceptodds
Under review as a conference paper at ICLR 2027

The Spatial Harness: From Fallible Visual Tools to Reliable Agentic Spatial Reasoning

Abstract

While vision-language models (VLMs) excel at semantic understanding, they struggle to locate evidence from videos or multi-view images, and express it as precise spatial measurements, such as object sizes and distances. Visual tools can supply these locations and measurements, yet their outputs remain fallible and may introduce plausible but incorrect observations into downstream reasoning. Reliable tool-assisted reasoning therefore requires verifying these observations before accepting them as evidence and seeking alternatives when they are rejected or insufficient. Building on this principle, we introduce **Spatial Harness**, which models visual tools as stochastic observation channels and formalizes evidence acquisition as an output-feedback control process. Implemented as a planner–executor–verifier loop, the planner is the acquisition controller, proposing primary and backup routes for the executor to instantiate, while verification and memory updates jointly provide observer-like interpretation of tool outputs. Updated memory guides replanning for evidence still missing, while routes retaining no evidence trigger backup execution. To train the roles for distinct responsibilities while reducing reward conflicts, we first optimize specialized teachers on offline role-local contexts containing successful and failed tool outputs. We consolidate their expertise into one role-conditioned agent through multi-teacher on-policy distillation, combining teacher feedback on student rollouts for capability acquisition with task-reward reinforcement learning for refinement. On ReVSI and Ego3D-Bench, the harness improves the accuracy of open-weight models and closed-source APIs, outperforming spatially specialized models. Post-training yields further gains by improving collaborations between agents and imperfect tools.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.