Accuracy, Architecture, or Method: What Determines XAI Quality in Vision Models?
Abstract
Explainable Artificial Intelligence (XAI) is an emerging topic which aims at making opaque AI models’ decisions transparent and understandable. Various methods have been proposed to explain deep learning models across different domains. What still remains an ongoing debate is the question on how the explanation quality of XAI methods should be evaluated. Therefore, numerous quantitative metrics have been suggested which are supposed to measure certain explanation properties such as localization, faithfulness, complexity and randomization. Previous studies, mostly using off-the-shelf pretrained backbones, have indicated that explanations might be influenced by three factors: the model architecture, the model performance, and the XAI method. In addition, the interplay between these three factors has not been investigated yet. Our study contributes to this research gap by determining how much these three factors actually impact the nominal explanation quality of plain image classifiers. We therefore train 26 different Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) from scratch, spanning model sizes from 3.4M to 210M parameters. In our study we use the same training parameters on a 50-class ImageNet subset with pixel-level object masks, and score five XAI methods with twelve evaluation metrics at eight checkpoints per model, covering accuracies from 16% to 91%. From our image classification scenario three results emerge: 1. Accuracy has the largest impact on localization. On the other hand, complexity, faithfulness and randomization metrics are mainly influenced by the XAI method and do not significantly improve as the model learns. 2. With ongoing training, perturbation-based XAI methods tend towards outperforming gradient-based methods on transformers regarding most metrics we measure. The opposite behaviour can be observed for CNNs. 3. Within a model architecture it are the normalization layers that have the highest impact on localization metrics. Simply increasing the number of model parameters, however, does not yield measurable improvements regarding any XAI metric.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.