acceptodds
Under review as a conference paper at ICLR 2027

They Know Where It Points, But Not What It Means: The Causal Grammar of Relational Visual Prompts

Abstract

Graphical cues can identify targets and express relations. We test whether vision-language models use an arrow to attribute speech after locating its endpoint. Across 450 paired conditions from 150 scenes, five open VLMs locate endpoints 8.5–13.1 percentage points more accurately than they attribute speakers, even when given the arrow-to-speaker rule. Adding the cue changes candidate scores, and decoding these changes improves speaker attribution by 8.5–15.7 points. Complete three-target operator orbits recover relation-sensitive responses when raw answers remain unchanged, outperforming matched shams on a controlled benchmark and authored comics. A second directed relation reproduces the behavioral gap and cue-specific recovery. Activation replacement finds an early visual effect that fades as an answer-position effect emerges. Under a reversed semantic codebook, predictions remain more closely aligned with the physical operator than with the instructed mapping. These results distinguish reference grounding, sensitivity to the operator, and use of the relation to select an answer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.