acceptodds
Under review as a conference paper at ICLR 2027

THE READOUT DOES NOT ROTATE: REFERENCE-FRAME BINDING FAILURE IN VISION-LANGUAGE MODELS

Abstract

Vision-language models frequently err in spatial perception when direction is defined by a reference object's heading. To investigate the cause, we fix the prompt, the target, and the camera position in controlled synthetic scenes, reverse only the reference object's heading, and examine eight open-source models stage by stage across perception, transmission, and answering. We find that both the reference-object heading and the camera information are linearly decodable and are carried to the answer token, yet the heading information is not used, and the model preferentially relies on the camera information instead. We then attempt a targeted fine-tuning remedy, but find that the model only memorizes the third-party headings it has seen and does not generalize what it has learned. Taken together, we conclude that the failure of perception in a third-party coordinate frame occurs in the mapping from representations to answers, and that fine-tuning does not repair this defect. We propose a set of metrics as a quantitative standard for this ability, which can test whether a proposed remedy genuinely establishes reference-frame binding and thereby support the evaluation of future work.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.