acceptodds
Under review as a conference paper at ICLR 2027

Anchor3D: Do 3D LLM Really Use 3D Positions?

Abstract

Recent 3D large language models (3D LLMs) inject explicit geometry into a language backbone through 3D positional embeddings, voxel-aligned features, or specialized rotary encodings, and report steady gains on 3D question answering and grounding benchmarks. We ask whether these models actually use the geometry they are given. Across five representative 3D LLMs, including a model with a dedicated 3D rotary encoding, we find that randomly permuting the 3D positions assigned to visual tokens leaves performance essentially unchanged across the benchmarks we evaluate, and so does setting them all to zero. Simple MLP probes show that, the LLM's hidden states can still carry accurate relative 3D offsets between tokens, so the geometry is encoded but not consulted: current benchmarks can be solved from appearance and language cues, and a purely autoregressive decoder is never forced to bind the entities named in a question to their 3D locations. Motivated by this, we propose to make question-conditioned localization an explicit bottleneck. Our model first detects and segments the objects named in the question, then selects one mask per named object with a question-conditioned matcher, and passes only the patch tokens covered by these masks, together with their 3D positions, to the LLM for metric reasoning. Trained on ScanQA, SQA3D, and R3D-Bench-style metric questions built on ScanNet, the resulting model, Anchor3D, becomes sensitive to positional perturbation, matches GPT-5.5 and Gemini 3.1 Pro on R3D-Bench with 4B parameters, and improves ScanQA and SQA3D over the same pipeline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.