acceptodds
Under review as a conference paper at ICLR 2027

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

Abstract

While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose GeoPID, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. GeoPID decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.