LESS THAN MEETS THE EYE: ACTION-SET IDENTIFIABILITY AND PRINCIPLED ABSTENTION FOR 3D PHYSICAL GEOMETRY
Abstract
A depth prediction can match its reference measurement and still describe a surface that no hand would touch. A corridor, its reflection in a planar mirror, and a photograph of that corridor on a screen can produce one image, while a hand stops in three places. Existing work addresses such scenes by enlarging what a predictor reports, through multi-layer formulations, detectors for mirrors and glass, and calibrated uncertainty. Each such quantity, however, describes a prediction or a predictor at a fixed viewpoint. None of them registers that the observation carries no evidence, because one image fits every explanation. We therefore ask a different question. Given a bounded set of camera motions, can an observer recover the physical explanation at all? Our approach, Intervene3D, calls two explanations separable to the extent that they predict different observations after the camera moves. A scene counts as identifiable once some available motion separates every remaining explanation, and the system abstains otherwise. A maximin rule then selects the motion separating the weakest contending pair. The experiments we conducted begin on a benchmark whose explanations are pixel-identical at the reference view, where abstention lowers false physical certainty from to . Maximin selection there reaches a committed accuracy of across four mechanisms, whereas a belief-weighted objective reaches on mirrors. Turning to real photographs, our experiments on seven families reveal that all of them fail on to of images, while their confidence stays below chance of predicting these failures. Choosing the viewpoint that some hypothesis explains best then lowers ambiguous-region error by on all thirteen checkpoints. Moreover, a gate over frozen encoder features reaches out of fold and lowers the failure rate from to at coverage. Asking what the available views can settle therefore precedes any answer about the physical geometry of the scene.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.