acceptodds
Under review as a conference paper at ICLR 2027

Tracing Perceptual Transfer in Vision-Language Models via Model Merging

Abstract

Despite recent progress in visual mathematical reasoning, perception remains a major bottleneck for Vision-Language Models (VLMs). However, it remains unclear where perceptual transfer occurs within the visual hierarchy of VLMs and how its representational and functional effects vary across layers. To investigate this question, we use model merging as a tool and apply it to the perception modules of VLMs, while keeping the language model frozen, allowing us to understand how perceptual transfer occurs across layers. Across multiple VLMs, this intervention yields gains on several visual mathematical reasoning benchmarks (MathVista, MathVerse, MathVision), providing evidence that perceptual capabilities can be transferred without modifying the LLM. To understand where these gains arise, we further analyze perceptual transfer across the visual hierarchy. Layer-wise probing shows that primitive geometric attributes become linearly accessible relatively early in the visual hierarchy, whereas relation-sensitive information becomes more prominent in intermediate layers. We further perform backward and forward layer replacement to examine the functional effects of perceptual transfer. Removing merged parameters from intermediate layers causes the largest performance degradation, while inserting them into the base model yields the largest performance gains in the same layers. Together, these complementary analyses reveal a non-uniform layer-wise organization of perceptual transfer, with relation-sensitive information and its strongest functional effects concentrated in middle layers. These findings provide a layer-wise account of how perceptual knowledge is represented and transferred within VLMs, offering insights for targeted perception-side adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.