acceptodds
Under review as a conference paper at ICLR 2027

HiSuper-VLM: Multi-Scale Scene Understanding via Hierarchical Superpoints

Abstract

3D Vision-Language Models (3D-VLMs) have shown strong potential for 3D scene understanding. However, scaling to larger, more complex scenes challenges current 3D-VLMs in both efficiency and accuracy. This is largely due to their use of variable-length scene tokens that grow with scene complexity and the decoupling of scene representation and understanding, which limits VLM's adaptive reasoning over scenes with different scales. To address this, we propose HiSuper-VLM, a 3D-VLM framework leveraging hierarchical superpoints to construct compact yet expressive representations, providing a foundation for VLM understanding across multiple scales and hierarchical levels while tightly coupling representation with reasoning. Specifically, for representation, we construct hierarchical superpoints, using fixed-length large-scale superpoints as initial visual tokens for VLMs while retaining fine-scale superpoints for information completeness. For understanding, we perform multi-level queries across VLM layers, from coarse to fine, enhancing reasoning by refining 3D representations during the process. Experimental results demonstrate that HiSuper-VLM significantly outperforms state-of-the-art methods on multiple indoor and outdoor 3D scene understanding benchmarks. Code will be publicly released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.