Same Geometry, Different Scores: Auditing Coordinate and Scale Dependence in Geometric Video Evaluation
Abstract
Video generation benchmarks increasingly score geometry: they reconstruct camera paths, depth, and object motion from generated frames and score the reconstruction. The axes and scale of a reconstruction are arbitrary conventions, so rotating the axes or rescaling the scene describes the same scene, and a geometry score should not change. We test whether it does: we rotate or rescale saved reconstructions, compare the scores, and trace every change to the code that causes it; for WorldScore we check instead that a reversed camera path scores lower. The audited component of each benchmark fails its test. On simulator ground truth, the range of PDI-Bench's motion score across 256 sampled rotations has a median of 80.7% of the unrotated score (objects with a nonzero score), because the metric takes medians separately along each axis. Rescaling the scene by 10 changes the final MBench scores of 72% to 83% of 3,218 videos and doubles the gap between two generators, because its frame-pair selector adds a distance to an angle; switching the reconstruction model can, through scale alone, reverse the direction of a score change. In controlled translation tests, WorldScore gives full marks to a camera that moves the opposite way. On Dynamic Replica, a comparison of two pipeline versions reverses once both are scored on the same observations. In the three benchmark components, each failure comes from a few lines of code, and minimal code changes pass the tested relations. These tests are cheap, need no new videos, and should be run before geometric scores are used to compare models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.