Vision-Language Models Can Report Which Image's Visual-Token Activations Were Perturbed
Abstract
Can a vision-language model (VLM) report which of two images had its internal visual-token activations perturbed while the visible inputs themselves remain unchanged? We study this question in three instruction-tuned VLM checkpoints by perturbing decoder-side activations at candidate-associated visual-token positions and asking the model which candidate was affected. Because successful localization could simply reflect damaged visual information, we separately test whether measured visual functions remain intact and use controls to probe simpler explanations. At operating points satisfying both the category criterion and a secondary counting guard, localization rises from \(0.499\) to \(0.688\) in Qwen2.5-VL and from \(0.505\) to \(0.656\) in Qwen3-VL, with both effects surviving whole-window multiplicity correction. Idefics3 shows a smaller but statistically detectable effect. Norm-restored and crossover controls show that activation-norm inflation and fixed display-side preference are not sufficient explanations, while matched-quality and visible-corruption comparisons bound stronger perturbation-specific interpretations. Finally, in Qwen3-VL, perturbation identity is near-ceiling linearly decodable at both \(p=.30\) and \(p=.60\), yet the fitted probe directions have sharply different steering efficacy: little targeted steering at \(p=.30\) but material bidirectional steering at \(p=.60\), where the counting guard no longer holds. We call this phenomenon perturbation reportability: candidate-specific internal perturbation information can become behaviorally reportable, while linear availability and steering efficacy remain distinct under the tested intervention procedure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.