acceptodds
Under review as a conference paper at ICLR 2027

What a 3D Reconstruction Benchmark Can and Cannot Tell About Geometric Foundation Models

Abstract

Benchmarks for 3D geometric foundation models rank models by surface metrics such as Chamfer distance, without testing whether those metrics can tell a reconstruction of the right scene from one of the wrong scene. This paper tests it for fourteen models on six domains by scoring two substitutes through the unmodified pipeline: another scene's ground truth, planted on the target's cameras, and the target's own ground truth flattened onto a plane. Where the ground has relief, the wrong scene is rejected on 98% of cells. On low-relief ground, rendered lunar terrain and rover imagery of Mount Etna and Saharan sand alike, the flattened ground truth outscores every model on all 21 windows, and on about half of them even when every submission receives the same similarity fit. A closed-form bound explains why: a plane's Chamfer cannot exceed the target's mean deviation from its own plane, so the plane wins wherever the best model errs by more. The wrong scene also outscores every model on most low-relief cells, but on the real rover data only through placement, which the shared similarity fit removes. The ground-level tables here and those the field publishes are too small to separate their leading models once every pair is tested. The controls and the evaluation code will be released, with a checklist for benchmarks on low-relief terrain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.