Question Tokens Relay an Answer-Level Summary in Vision-Language Models
Abstract
Activation patching asks how vision-language models (VLMs) use visual information by copying hidden states from a donor input into a recipient input during inference. However, changing the donor image usually changes both the image content and the correct answer, confounding the two. To disentangle these factors, we introduce a controlled activation-patching setup that holds the question fixed while varying whether the donor shares the recipient's answer. We construct 538 quartets from VQAv2, each pairing one question with four images, two supporting each of two answers, e.g., two with ”yes” and two with ”no”. Within each quartet, we patch the recipient with question-token states from either a same-answer or different-answer donor at a given layer, while matching per-token norms. In all seven checkpoints we study, different-answer donors pull the recipient toward the donor's answer; in six, same-answer donors change the full output distribution far less. Yet same-answer donors are not neutral: in four checkpoints, a recipient moves one-third to two-thirds of the way toward their donors' own answer margins. Later layers thus read an answer-level summary from the question tokens, namely which answer the image supports and how strongly, and a layer sweep in three models places this summary in a narrow band of middle layers. In each of the three VLMs, a learned rank-16 component preserves the donor effect when retained and suppresses it when removed, more strongly than magnitude-matched controls. The donor effect is weakest when the intervention layer lies beyond this window, so patching studies should control both the donor's answer and the layer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.