Do 3D Large Language Models Really Use 3D Evidence?
Abstract
Recent 3D Large Language Models (3D-LLMs) have shown strong performance on scene question answering with extensive post-training. Yet, existing benchmarks rarely evaluate how 3D-LLMs respond to referential ambiguity, insufficient scene evidence, or false presuppositions, obscuring whether their responses are grounded in 3D evidence. In this work, we introduce UNICORNS, a benchmark for evaluating the reliability of 3D-LLM responses across four evidence conditions. UNICORNS contains 11,387 questions across 557 ScanNet scans, covering eight challenge categories alongside answerable controls. A response passes only if it satisfies the required response behavior and contains no scene claim unsupported by the available evidence. Evaluation of six open-sourced 3D-LLMs reveals a pronounced performance gap: even the strongest zero-shot model achieves 61.22% accuracy on answerable controls, but only 40.94% across the eight challenge categories. To address this gap, we further propose HORN, an obligation-aware post-training method combining group-relative task rewards, leaf-conditioned obligation rewards, and supervised control replay. Extensive experiments demonstrate that HORN reaches 75.1±1.0% across the challenge categories and 77.2±0.8% overall, outperforming all baselines without degrading answerable-control accuracy. Our findings enable a systematic evaluation and improvement of evidence-grounded 3D reasoning for 3D-LLMs(code).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.