acceptodds
Under review as a conference paper at ICLR 2027

RIGHT PREDICTIONS, MISLEADING EXPLANATIONS: ON THE VULNERABILITY OF VISION–LANGUAGE MODEL EXPLANATIONS

Abstract

Explanation mechanisms are increasingly used to support transparency and trust in vision–language models (VLMs), particularly in settings where model decisions require human oversight. However, the robustness of these explanations remains insufficiently understood. In this work, we investigate whether explanation heatmaps in VLMs, particularly CLIP-based models, remain reliable under adversarial conditions. We show that explanation maps can be systematically manipulated while preserving the model’s original prediction, revealing a disconnect between predictive behavior and the evidence the explanation presents. To study this vulnerability, we introduce X-Shift, a white-box attack that perturbs the input image to redirect the explanation heatmap of the predicted class onto the region that the explainer associates with a second, attacker-chosen label, without altering the predicted output. Unlike conventional adversarial attacks that aim to induce misclassification, X-Shift specifically targets the integrity of the explanation process itself. The attack operates without modifying model parameters. We evaluate the proposed approach on MS-COCO, Flickr30k and ImageNet-1k with three CLIP backbones, demonstrating that the explanation is relocated onto the target region under small perturbations on image–backbone pairs while the prediction is preserved, with every result compared against random noise of the same budget. Our findings highlight a fundamental limitation of current explanation mechanisms in VLMs and raise concerns about their use as reliable indicators of model trustworthiness in high-impact applications.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.