Correspondence, Not Similarity: Gating Multi-View Evidence in Driving Vision-Language Models
Abstract
Driving vision-language models are given several cameras, and a growing number of evaluations conclude that a model used the extra view. The usual evidence is a matched negative: show the model a comparable image containing no target object, and check that its answer returns to baseline. We show that this control certifies less than it is taken to certify. Holding scene, timestamp, ego pose and ground truth fixed across 610 Paired Exposure Units built from nuScenes, we replace the true side camera with one from another vehicle's world, matched on city, hour, weather and road-user status. Across five vision-language models the true and the substituted view prove equally persuasive, and two models find the substituted view more persuasive. The distinction is easy to draw from pixels, where an off-the-shelf matcher separates the two at 0.973 AUROC. We benchmark ten verification methods and find that those measuring correspondence keep the evaluation valid at every difficulty, while those measuring similarity fail regardless of supervision. Gating on correspondence needs no extra generation. We release the benchmark, the pools, and all 122,000 responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.