acceptodds
Under review as a conference paper at ICLR 2027

FrameBind: Directional Grounding for Allocentric Spatial Reasoning with VLMs

Abstract

Vision-Language Models (VLMs) have made rapid progress on spatial reasoning benchmarks, but it remains unclear whether these gains reflect grounded geometric understanding or semantic shortcuts. We study allocentric relative-direction tasks, where the models are prompted to determine the direction a target object is with respect to reference object or to determine the object in a given direction. On controlled scenes we find a pronounced answer-format asymmetry: models appear to be more accurate at selecting an object than at naming the direction, even when the image and underlying relation are held fixed. We trace this gap to a strong camera-centric bias; in our experiments, only 3.6% of a representative model's direction answers match the requested reference frame, while 95% match the camera frame. Component diagnostics show that VLMs localize objects reliably but cannot recover a reference object's orientation or extend partial orientation cues into a usable frame. Only a complete labeled coordinate display can restore performance, likely because it reduces the task to image-plane localization. Motivated by this diagnosis, we introduce FrameBind, a training-free, tool-assisted pipeline that overlays a reference-centered coordinate frame on the original RGB image rather than replacing it with a symbolic abstraction. Additionally, we introduce an Object-Map Inversion that converts direction queries into sparse, direction-to-object maps to exploit the object-answer preference. Across four benchmarks, FrameBind improves both answer formats over its Qwen3-VL-8B backbone, while object-map inversion adds further gains on controlled scenes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.