acceptodds
Under review as a conference paper at ICLR 2027

GeomFunnel: Coordinate-Anchored Sparse Voxel Funneling for Metric-Aware 3D-VLMs

Abstract

Modern 3D vision-language models increasingly compress rich 3D scene evidence into compact representations for efficient reasoning, inevitably losing some information in the process. Our design addresses the risk that semantic scene understanding and relative spatial relationships can remain accessible after compression while the fine-grained metric geometry required for precise physical measurement becomes harder to recover. To address this gap, we introduce GeomFunnel, which constructs a coordinate-grounded sparse physical field that anchors multi-view scene evidence in a shared metric 3D space and preserves it as an explicit physical substrate. Specifically, its representation-decoupled reasoning uses compact tokens for semantic and relative spatial reasoning while retaining direct access to physical geometry for queries that require precise metric quantities. Within the metric pathway, language-guided physical grounding connects the referred scene elements to their relevant physical supports, allowing the requested quantities to be computed from explicit geometry rather than recovered from compressed tokens. GeomFunnel achieves a new state-of-the-art average score on VSI-Bench while remaining competitive on general 3D vision-language benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.