acceptodds
Under review as a conference paper at ICLR 2027

H3D-LVR: Hierarchical 3D Latent Visual Reasoning for Multi-View VLMs

Abstract

Vision-language models have achieved strong progress in visual perception and instruction following, yet multi-view 3D spatial reasoning remains challenging. Although textual chain-of-thought can expose intermediate reasoning steps, it operates at the output level and does not directly constrain the model to form spatially correct intermediate hidden states. Existing 3D or latent visual reasoning methods provide spatial supervision, but typically rely on a single-level latent representation, which is insufficient to capture the multi-granular information required for multi-view reasoning. We propose H3D-LVR, a hierarchical 3D latent visual reasoning framework that organizes a VLM's reasoning trajectory around three generated latent states: scene layout, key objects, and input views. During SFT, VGGT-derived layout features and matched object/view features supervise the corresponding spans through reconstruction losses. The corresponding features also fill latent-pad inputs under teacher forcing, but no such features are provided at inference. We further adapt answer-level GRPO to this hierarchical latent format with hidden-state replay and full-span loss masking, allowing reinforcement learning to improve answer behavior without corrupting the learned latent-token semantics. During inference, H3D-LVR uses only the input images and prompt; it does not run VGGT, receive precomputed layout/object/view features or annotations, or invoke external 3D tools. Experiments on MindCube-Tiny, SPAR-Bench, and MMSI-Bench demonstrate that H3D-LVR achieves substantial performance gains, strongly validating its effectiveness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.