acceptodds
Under review as a conference paper at ICLR 2027

What You See Is Not What Matters: Attacking VLM Explanation Faithfulness

Abstract

Visual explanations are increasingly used to understand the decisions of Vision-Language Models (VLMs), yet their robustness under adversarial perturbations remains poorly understood. Existing studies mainly focus on saliency changes, but a substantial explanation shift does not necessarily imply unfaithfulness. We investigate whether prediction-preserving perturbations can redirect saliency toward regions that provide weaker support for the model's prediction. To support this evaluation, we propose the Faithfulness-Oriented Adversarial (FOA) Attack, which uses explanation faithfulness as an optimization signal to stress-test highlighted regions. By detaching saliency and mask generation, FOA avoids backpropagation through the explainer while enabling faithfulness-oriented updates. Across five VLM explanation methods and three datasets, our evaluation shows that faithfulness can be substantially degraded while preserving the original prediction, with FOA consistently revealing stronger vulnerabilities than X-Shift, a state-of-the-art attack on VLM explanations that primarily targets saliency shifts rather than explanation faithfulness. FOA perturbations also transfer effectively across explainers, suggesting that these vulnerabilities extend across explanation methods. These results demonstrate that highlighted regions can become less representative of the evidence supporting a prediction even when the prediction is unchanged. Our findings show that explanation robustness should be assessed not only by where explanations shift, but also by whether the resulting explanations remain faithful.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.