Pose Shortcuts and Frame Reduction: When Do Canonical Views Help Spatial Verification?
Abstract
Rendering a relation-conditioned canonical view of a reconstructed scene is a natural way to make a spatial relation visually verifiable. We ask when such a view actually helps a verifier, using fixed-reference relations in 3D Gaussian splatting reconstructions of ScanNet rooms, and find three answers. First, a canonical camera is a function of the queried geometry: for left–right and front–back the camera rotation alone determines the room-frame label in 100% of our views, so a trained verifier that receives the pose need not read the image. A pose and-mask probe without RGB reaches 99.6% and 91.1% on canonical views of ten previously unused rooms and inverts to 0.2% and 9.6% when only the rotation is replaced by the mirrored camera's; the ResNet verifier itself inverts from 66.2% to 33.9% under the same mirrored rotation, so its front–back training-by-view interaction ( points) is a pose effect. Second, an analytic ceiling shows that a mask-conditioned verifier built on canonical views cannot beat the noisy geometry that placed its camera except when the noise reverses the pair direction, whereas gravity-aligned cameras keep the up–down frame conversion exact; with 1 m center noise and correctly marked objects, GPT-4o verifying up–down in such cameras beats the noisy geometry by 9 points. Third, for vision–language models the value of a canonical view is frame reduction: asking the question in the image's own frame and converting the answer with the known camera axes makes canonical views beat default views on left–right for all four models we test, by 6 to 23 points (90.1% for GPT-4o), while the room-frame interface stays at 50–58% in every view. Left–right pairs a model tells apart in the image frame are almost always right; depending on the model, the canonical gain comes from fewer undecided pairs, an exact frame conversion, or both. Canonical views help when they turn a frame-composition problem into a perception problem the verifier can solve; reported gains must be checked against pose shortcuts and geometry ceilings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.