acceptodds
Under review as a conference paper at ICLR 2027

View-AID: Multi-View Spatial Reasoning via Auxiliary-Image Distillation

Abstract

Multi-view spatial reasoning requires multimodal large language models to reconcile fragmented observations into a coherent spatial state. Direct textual chain-of-thought (CoT) construction serializes this global state into local relational statements, leaving cross-step spatial consistency implicit. Conditioning CoT generation on a reference answer further specifies the endpoint without constraining the intermediate spatial process. We introduce View-AID, which improves spatial supervision at its source by externalizing task-relevant spatial hypotheses as auxiliary images before translating them into textual CoT. This visual-first construction enriches reasoning supervision while preserving a standard inference interface that uses only the original views and question. Confidence-Guided Policy Optimization (CGPO) then complements binary outcome rewards with trace-conditioned answer confidence. On Qwen3-VL-4B, View-AID improves MindCube-Tiny and MMSI-Bench by 48.4 and 9.4 percentage points, respectively, and exceeds the strongest listed spatial baseline by 4.2 points in Overall accuracy. The gains transfer to disjoint spatial reasoning tasks and scale effectively to a larger backbone, supporting visual organization before textualization as an effective approach to multi-view spatial reasoning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.