acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language Models Cannot Tell What Moves Without Knowing How They Move

Abstract

Vision-language models for driving are asked whether a nearby vehicle is moving or stationary, and the usual remedy for their errors is to show them more frames. We find that more frames do not help. A BLIP-2 model finetuned on nuScenes scores 62.1% from one frame and 62.9% with two past frames fused through gated cross-attention, and three open vision-language models change by−3.8 to +2.7 points when given three frames instead of one. We prove that this failure concerns information and not architecture. The motion state of an object is not identifiable from its positions in the camera frame, because that frame mixes the motion of the object with the apparent motion that the observer causes, and it becomes identifiable once the relative pose of the camera between frames is known. On nuScenes, object displacement separates the two classes with an AUC of 0.557 in the ego frame and 0.935 after ego-motion compensation, and one token that carries the compensated displacement raises the finetuned model to 88.2 ±0.3%, or 88.1 ±0.4% with the image blacked out. How the quantity enters the model matters as much as whether it is present: the same twelve numbers give 66.2% through a subtractive structure that cannot represent rotation and 76.2% with auxiliary supervision of that structure, whereas a module that builds the exact group action in reaches 88.1 ±0.2%, and stating the compensated displacement in the prompt of an open model without training adds between 0 and 10 points. A model on a moving platform needs its own motion as an input and the compensation as a structure, not a longer clip.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.