Towards Efficient and Effective Scene Visual Representation for 3D Spatial Scene Understanding
Abstract
Despite recent advances in multimodal large language models (MLLMs), 3D spatial reasoning remains challenging because spatial evidence is often redundant, occluded, or fragmented across densely sampling views. Existing video-based 3D MLLMs typically adapt video MLLMs to 3D tasks while retaining dense visual tokens, resulting in heavy inference costs. Recent post-hoc token compression methods can alleviate these costs but inevitably compromise performance. In this paper, we propose a unified framework that integrates aggressive visual token compression and multi-teacher distillation into a single stage of 3D task fine-tuning, dubbed SAVR. Specifically, we compress visual tokens to a prescribed token budget before they enter the LLM, enabling the model to adapt directly to compact visual tokens. To preserve informative visual cues under aggressive compression, we distill complementary priors from diverse vision foundation models into visual token representations at the final LLM layer. A learnable adaptive weighting mechanism further modulates the contributions of individual teachers to mitigate potentially noisy supervision and facilitate selective knowledge transfer. By coupling token compression with task adaptation and complementary expert distillation, our method leads to compact, task-relevant visual representations without requiring a separate post-hoc compression stage and maintains relatively vanilla performance. Extensive experiments across multiple 3D vision and spatial reasoning benchmarks demonstrate SoTA token efficiency with negligible performance degradation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.