ESCBench: Benchmarking Context-Sensitive Spatial Deixis in Vision-Language Models
Abstract
Evaluating spatial demonstratives requires asking whether a response to context retains its direction and distance profile across observations of the same situation. Recognition accuracy, prompt stability, and aggregate human resemblance each leave this question unresolved. We introduce ESCBench, a diagnostic benchmark of English this and that that crosses changes in depicted context with changes in observation under a fixed speaker task. Its 576 tabletop scenes vary distance, tool configuration, interlocutor layout, shape, and color, yielding 2,304 images from four viewpoints. Linked comparison identities support tool and layout contrasts within each view; seven protocols connect these contrasts to attribute recognition and elicitation sensitivity. Across 22 vision–language model deployments, the same tool contrast can reverse across views despite high overall recognition and low overall task-form sensitivity. Similar external-human distribution scores can also accompany a large positive tool response or a near-zero one, while pooling views can yield a closer match than every individual view. These findings distinguish contextual responses that agree across observations from observation-dependent responses and aggregate resemblance produced by mixing them. ESCBench makes these distinctions directly testable through matched comparisons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.