acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language Models Can Report Which Image's Visual-Token Activations Were Perturbed

Abstract

Can a vision-language model (VLM) report which of two images had its internal visual-token activations perturbed while the visible inputs themselves remain unchanged? We study this question in three instruction-tuned VLM checkpoints by perturbing decoder-side activations at candidate-associated visual-token positions and asking the model which candidate was affected. Because successful localization could simply reflect damaged visual information, we separately test whether measured visual functions remain intact and use controls to probe simpler explanations. At operating points satisfying both the category criterion and a secondary counting guard, localization rises from \(0.499\) to \(0.688\) in Qwen2.5-VL and from \(0.505\) to \(0.656\) in Qwen3-VL, with both effects surviving whole-window multiplicity correction. Idefics3 shows a smaller but statistically detectable effect. Norm-restored and crossover controls show that activation-norm inflation and fixed display-side preference are not sufficient explanations, while matched-quality and visible-corruption comparisons bound stronger perturbation-specific interpretations. Finally, in Qwen3-VL, perturbation identity is near-ceiling linearly decodable at both \(p=.30\) and \(p=.60\), yet the fitted probe directions have sharply different steering efficacy: little targeted steering at \(p=.30\) but material bidirectional steering at \(p=.60\), where the counting guard no longer holds. We call this phenomenon perturbation reportability: candidate-specific internal perturbation information can become behaviorally reportable, while linear availability and steering efficacy remain distinct under the tested intervention procedure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.