View-AID: Multi-View Spatial Reasoning via Auxiliary-Image Distillation
Abstract
Multi-view spatial reasoning requires multimodal large language models to reconcile fragmented observations into a coherent spatial state. Direct textual chain-of-thought (CoT) construction serializes this global state into local relational statements, leaving cross-step spatial consistency implicit. Conditioning CoT generation on a reference answer further specifies the endpoint without constraining the intermediate spatial process. We introduce View-AID, which improves spatial supervision at its source by externalizing task-relevant spatial hypotheses as auxiliary images before translating them into textual CoT. This visual-first construction enriches reasoning supervision while preserving a standard inference interface that uses only the original views and question. Confidence-Guided Policy Optimization (CGPO) then complements binary outcome rewards with trace-conditioned answer confidence. On Qwen3-VL-4B, View-AID improves MindCube-Tiny and MMSI-Bench by 48.4 and 9.4 percentage points, respectively, and exceeds the strongest listed spatial baseline by 4.2 points in Overall accuracy. The gains transfer to disjoint spatial reasoning tasks and scale effectively to a larger backbone, supporting visual organization before textualization as an effective approach to multi-view spatial reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.