GeomFunnel: Coordinate-Anchored Sparse Voxel Funneling for Metric-Aware 3D-VLMs
Abstract
Modern 3D vision-language models increasingly compress rich 3D scene evidence into compact representations for efficient reasoning, inevitably losing some information in the process. Our design addresses the risk that semantic scene understanding and relative spatial relationships can remain accessible after compression while the fine-grained metric geometry required for precise physical measurement becomes harder to recover. To address this gap, we introduce GeomFunnel, which constructs a coordinate-grounded sparse physical field that anchors multi-view scene evidence in a shared metric 3D space and preserves it as an explicit physical substrate. Specifically, its representation-decoupled reasoning uses compact tokens for semantic and relative spatial reasoning while retaining direct access to physical geometry for queries that require precise metric quantities. Within the metric pathway, language-guided physical grounding connects the referred scene elements to their relevant physical supports, allowing the requested quantities to be computed from explicit geometry rather than recovered from compressed tokens. GeomFunnel achieves a new state-of-the-art average score on VSI-Bench while remaining competitive on general 3D vision-language benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.