acceptodds
Under review as a conference paper at ICLR 2027

Do 3D Large Language Models Really Use 3D Evidence?

Abstract

Recent 3D Large Language Models (3D-LLMs) have shown strong performance on scene question answering with extensive post-training. Yet, existing benchmarks rarely evaluate how 3D-LLMs respond to referential ambiguity, insufficient scene evidence, or false presuppositions, obscuring whether their responses are grounded in 3D evidence. In this work, we introduce UNICORNS, a benchmark for evaluating the reliability of 3D-LLM responses across four evidence conditions. UNICORNS contains 11,387 questions across 557 ScanNet scans, covering eight challenge categories alongside answerable controls. A response passes only if it satisfies the required response behavior and contains no scene claim unsupported by the available evidence. Evaluation of six open-sourced 3D-LLMs reveals a pronounced performance gap: even the strongest zero-shot model achieves 61.22% accuracy on answerable controls, but only 40.94% across the eight challenge categories. To address this gap, we further propose HORN, an obligation-aware post-training method combining group-relative task rewards, leaf-conditioned obligation rewards, and supervised control replay. Extensive experiments demonstrate that HORN reaches 75.1±1.0% across the challenge categories and 77.2±0.8% overall, outperforming all baselines without degrading answerable-control accuracy. Our findings enable a systematic evaluation and improvement of evidence-grounded 3D reasoning for 3D-LLMs(code).

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.