acceptodds
Under review as a conference paper at ICLR 2027

Point clouds as visual tokens for pretrained multimodal large language models

Abstract

Multimodal large language models (MLLMs) for 3D scene understanding typically rely on multi-view 2D images or dedicated 3D encoders. However, the former incurs heavy visual computation and represents metric geometry indirectly, whereas the latter requires separate feature extractors and alignment with language under limited paired 3D-text supervision. To bridge this gap, we present a compact, geometry-informed 3D tokenizer that directly reuses a pretrained MLLM's visual encoder and connector. Our tokenizer pools volumetric patches into a bounded set of 3D regions, projecting their descriptors into visual tokens via a single linear layer. We use 3D rotary position embeddings (3D-RoPE) to make visual attention sensitive to relative positions between region centroids, alongside spatially constrained spectral ordering to sequence contextualized features for the causal decoder based on geometric proximity and visual similarity. Training proceeds in two stages: first aligning region features with RGB-D teacher targets via cross-partition distillation using shared point membership, then fine-tuning response generation with attention supervision on annotated referents. Experiments across 3D question answering and dense captioning benchmarks demonstrate strong performance with only 3.6-4.0M tokenizer parameters across two pretrained backbones.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.