acceptodds
Under review as a conference paper at ICLR 2027

Remembering isn't Updating: Motion-Guided Spatial Reasoning in Vision-Language Models

Abstract

Spatial reasoning requires understanding spatial relationships and keeping them consistent across changing viewpoints. Yet retaining past observations does not ensure that a vision-language model can interpret them in the reference frame required by a question. We propose FrameShift, a geometry-guided method for spatial reasoning that represents camera motion in both the initial and final reference frames. This dual-reference representation preserves complementary spatial information about the same trajectory. A learned readout uses the question to combine evidence from the two reference frames. FrameShift then combines question relevance with geometric confidence to control the contribution of spatial evidence. The resulting spatial tokens are integrated with visual and language representations for answer generation. We further introduce geometric supervision and consistency objectives that connect continuous spatial information with language predictions. FrameShift improves overall accuracy on SAW-Bench by 7.19 percentage points and relative-direction accuracy by 22.30 points over the base model. These results highlight the value of representing the same trajectory in complementary reference frames. They support using motion and question context together to interpret spatial relationships in the appropriate coordinate system as the viewpoint changes. Code is available at https://anonymous.4open.science/r/FrameShift-0368.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.