acceptodds
Under review as a conference paper at ICLR 2027

Language Conditioning Elicits Symbol Grounding in Visual Processing

Abstract

Grounding is a core ability for visual reasoning in embodied intelligence. Within multimodal models, prior work has identified mechanisms for grounding, such as binding and retrieval, in language models. We posit that language conditioning can also elicit grounding in visual processing. Specifically, we inject linguistic context into a frozen pretrained vision transformer (ViT) via cross-attention, shaping the representations of objects and their features to serve as potential grounders for linguistic entities. We evaluate this approach on visual question answering tasks across vision backbones and analyze the representations they encode. We examine the representations at global and local levels. First, geometric analysis shows that ViTs encode object representations with attribute-related geometric structure, with language conditioning reshaping this hierarchy and substructures according to the queried attribute. Further analysis on geometry across layers reveals two mechanisms for grounding: binding segregates scenes containing the referent in the middle layers; organization repartitions scenes into groups aligned with the queried attribute values for retrieval in the latter layers. Second, controlled experiments on embeddings demonstrate that linguistic context enhances attribute alignment of the referent and its queried attribute while suppressing those of non-referents and unqueried attributes; referent-selective modulation also holds for natural images. This study opens a new direction for studying symbol/language grounding, the one that treats perception itself as an active component of grounding and investigates how linguistic signals can shape its representation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.