GeoWorld-VLM: Geometry from World Models for Vision-Language Models
Abstract
Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between. We introduce GeoWorld-VLM, which transfers predictive representations from a frozen camera-conditioned video world model into a VLM's existing visual pathway to improve single-image spatial reasoning. From each static training image, a sampled camera trajectory and fixed scene-stability prompt produce viewpoint-conditioned teacher features without real multi-view observations. Post-projector feature alignment transfers this supervision, while a preservation anchor constrains changes to the visual interface consumed by the frozen language model. Only the image encoder, multimodal projector, and student alignment head are updated; the teacher and alignment heads are removed after training. Across two VLM backbones on What'sUp+VSR and EmbSpatial, GeoWorld-VLM improves over matched task-only fine-tuning by – percentage points and achieves the highest mean accuracy among the evaluated feature teachers. These results demonstrate that predictive world-model representations can improve single-image spatial reasoning at no extra inference cost.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.