Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
Abstract
Multiview 3D consistency metrics assess whether generated images depict a single coherent, static 3D scene (eg, MEt3R scores random noise images to be near perfectly 3D consistent). We find that neural metrics, including MEt3R, can assign high consistency to inconsistent image sets and disagree substantially with human judgments on novel view synthesis (NVS) outputs. To diagnose these failures, we use SysCON3D, a controlled robustness benchmark that systematically isolates 3D inconsistency. We decompose reconstruction-based neural metrics into backbone, residual, and aggregation components, yielding a parametric family of neural metrics for controlled analysis and metric improvement. Despite our best neural variant being more robust than MEt3R on SysCON3D, crucial failure modes persisted. Our analysis identifies the cause of failure: feed-forward reconstruction backbones, including VGGT, MASt3R, DUSt3R, and Fast3R, hallucinate geometry and cross-view support for unrelated scenes and random noise. Motivated by these limitations, we introduce COLMAP-based metrics that use verified matches, camera registration, dense support, and reconstruction failure as explicit evaluation signals, and achieve the strongest human alignment of all metrics in our experiments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.