acceptodds
Under review as a conference paper at ICLR 2027

Which output absorbs the conflict? Gauge-aware attribution for feed-forward 3D reconstruction

Abstract

Feed-forward 3D reconstruction models emit depth and cameras from a single pass, and we ask which of those outputs absorbs the disagreement when a scene's geometry conflicts with what the image looks like. The question is not well posed as it stands: for a model defined only up to scale a global depth rescale and a baseline rescale are observationally equivalent, so any "depth did x%, cameras did y%" split rests on a gauge convention the reader cannot see, and we fix one and compare only invariants. Nor is the natural statistic safe: the projection coefficient such decompositions invite is the product of directional alignment and magnitude, , rewarding a component for being large rather than for being right. Our instrument is a ray-preserving warp: every scene point moves along its ray through the reference camera, so that image is byte-identical at every dose while the other views change. The monocular cue is therefore held fixed by construction rather than by assumption, the re-rendered views stay geometrically self-consistent with a different relief instead of being degraded, and the conflict has a dose. Across 224 measurements, 104 (46%) score where explained variance is negative, worse than predicting no change, the worst at against ; an earlier stage of this analysis relied on such a coefficient, and we correct it here. Under explained variance VGGT's outputs jointly reproduce 51–94% of the conflict across four doses, declining monotonically in , though the registered form was not supported, while the relief it reports stays nearly inert (parallax-reliance slope , against 0.719 for RAFT triangulation with ground-truth poses at the same model-input resolution). DUSt3R and Fast3R reproduce none of it, in different ways, and the model whose outputs do reproduce it is the one with an explicit camera head, though it also differs in attention structure, so we state the architectural attribution as supported rather than established. Our identifiability claim is scoped to pointmap scale: a pre-committed rule declined to extend it to a joint disparity gauge over , and a normaliser-frozen control confirmed the decline was not an artifact.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.