Most of the Geometry Credited to Video Foundation Models Is a Ground Plane: Six Probe Choices, and the Small Effect That Survives Them
Abstract
A video foundation model is credited with 3D structure when depth decodes from its features, and a naive probe on driving footage reads R² near 0.85. A quadratic in the patch's own row and column reaches R²=0.844 on the same data; regress it out, and the probe that read 0.848 adds +0.122 over it. Most of that number is a ground plane. Image position is one of six choices a depth probe makes, with motion, noise timestep, block, temporal window and patch grid; none is consistently reported, and in an exploratory audit on ten open checkpoints each moves the verdict, by 1.6× to about 50× or a change of sign. We give a configuration that closes all six and test what survives. On 200 real driving clips, every choice controlled but the block, which is fixed the usual way, four of four predictive encoders carry depth beyond position and flow and zero of six generators do; selecting the block on held-out splits as well keeps the encoder count and lets one generator cross at +0.0188, 2.5 times below the weakest encoder. The classes separate by magnitude, not by kind: an effect worth 0.05–0.12 of variance, an order of magnitude below the headline numbers it replaces. It survives a dense optical-flow control, which narrows the margin by a third, and a VAE control, after which encoders keep 70% of their advantage. Representational convergence, often read as corroboration, points the other way: the converging models that align patch for patch share a subspace that transfers a depth probe between them at ceiling, yet position-residualized depth decodes from it in 0/5 models, so what independently trained video models agree on is where things sit in the frame. Harness, per-experiment results, preregistration and decision ledger are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.