Object-Token Influence in Vision–Language Models: Where It Lives and How It Adds Up
Abstract
Vision–Language Models (VLMs) extend language models with an image encoder, but how visual tokens covering objects operate inside the LM backbone remains less understood than in text-only settings. Prior work has typically examined different forms of object-token influence separately, without comparing them under a common depth statistic. We study object-token influence in seven adapter-style VLMs along three operationally distinct axes—erasure sensitivity, cross-instance transfer, and category readout. Placing all three on a common half-mass depth coordinate, we find the same order in every model: erasure sensitivity precedes cross-instance transfer, which in turn precedes category readout. The two intervention axes concentrate in shallow-to-mid layers, whereas the readout axis concentrates in mid-to-deep layers, yielding an intervention–readout gap. Whole-set object-token interventions exhibit supra-additive effects relative to a matched background contrast; this excess concentrates in shallow layers and is indistinguishable from zero at deeper layers. These findings identify candidate depths for targeted intervention and motivate visual-token pruning criteria that account for joint rather than purely tokenwise effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.