acceptodds
Under review as a conference paper at ICLR 2027

Who Acts on Whom? Probing the Binding of Objects to Roles in Vision Transformers

Abstract

Humans understand scenes not only by recognizing which objects are present but also by determining who acts on whom, that is, by assigning objects to event roles, such as agent and patient. Replicating this human capability could provide artificial intelligence with a foundation for representing and reasoning about relations between objects. Recent work shows that Vision Transformers (ViTs) learn to encode, in low-dimensional subspaces of their representations, whether two patches belong to the same object. However, it remains unclear how these models represent the binding of different objects to agent and patient roles. CLIP-style models with ViT image encoders are known to struggle to distinguish captions in which these roles are swapped, raising the question of whether role information is absent from their representations or present but insufficiently used. In this paper, we train a bilinear probe on ordered pairs of object patches and find that the binding of objects to roles is encoded in a *relation-binding subspace* of patch representations in CLIP-style models. Given the region of one object in a relation, the probe assigns high scores to the region of its partner (the agent given the patient, and vice versa). Ablating this subspace reduces image-text matching accuracy on captions with role swaps more than on any other type of caption modification, whereas ablating a random subspace of the same dimension has no such effect.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.