acceptodds
Under review as a conference paper at ICLR 2027

Inject3D: Re-Accessing Dense Geometric Features with Compact Tokens for Fine-Grained 3D Understanding

Abstract

Fine-grained 3D understanding requires access to local geometric evidence, yet existing 3D multimodal large language models typically compress dense point clouds into a small set of point tokens before language modeling. This compression can discard geometric details that subsequent LLM layers can no longer directly access, while increasing the point-token budget adds computational and memory overhead and complicates optimization. We introduce Inject3D, which feeds a compact set of spatially anchored point tokens into the LLM and retains high-resolution point-wise features as external geometric memory. At selected LLM layers, each anchor token retrieves features from its local 3D neighborhood and integrates them into its hidden state via cross-attention. To encourage effective use of the retrieved geometric evidence, we further introduce task-specific reinforcement post-training across captioning, visual question answering, and grounding. Experiments on object- and part-level tasks show that Inject3D effectively enhances fine-grained 3D understanding with a compact point-token sequence and substantially lower computational cost than increasing the input token budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.