acceptodds
Under review as a conference paper at ICLR 2027

The Hidden Evolution of Disguised Visual Context inside VLMs

Abstract

Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into consistent representations and interact with the text token space depends entirely on the integration architecture. Existing architectures typically either treat visual tokens as in-context prompts within the input sequence or inject them directly into the LLM's intermediate layers. A controlled comparison and understanding of how these architectural choices affect visual information and its internal transformation to integrate with the LLM remains underexplored. In this study, we provide a fair comparison by evaluating in-context and layer-wise injection Vision-Language Model (VLM) integration paradigms under identical training conditions on benchmarks spanning across single-image, multi-image, and video understanding. Our results show that in-context injection consistently outperforms layer-wise injection, with the largest gaps on OCR and video tasks, even though both allocate attention to visual tokens in similar patterns. This raises the question of what happens to visual information inside the LLM and how the LLM utilizes it. Investigating this, we uncover a hidden evolution of visual tokens, which enter the LLM as disguised visual context from a different representation space, lacking the structure of LLM representations, before being reshaped differently according to the integration paradigm. Our analysis reveals that this reshaping leads the visual representations in the LLM to capture fundamentally different frequency characteristics, which reflect whether they encode fine-grained local detail or coarse global structure. In-context injection progressively builds high-frequency, fine-grained detail, whereas layer-wise injection remains biased toward low-frequency features. We show that this evolution inside the LLM determines what visual features the VLM can capture and utilize effectively, such as the fine-grained detail required for text-centric tasks. It also determines whether visual representations converge toward the language space. Specifically, in-context visual tokens remain separate from text tokens through the middle layers and gradually merge with them by the final layers, whereas layer-wise visual tokens remain separated across layers. Since attention patterns are similar, the performance gap must instead come from the quality of the visual representations that each paradigm shapes at every layer. We confirm this through a causal intervention by suppressing the high-frequency content in the visual tokens inside the LLM, which degrades in-context injection across tasks, especially in text-centric ones, while layer-wise injection is far less affected. This shows that high-frequency content drives the advantage of in-context injection over layer-wise injection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.