acceptodds
Under review as a conference paper at ICLR 2027

ITXGNN: Learning Stable Explanations in Self-Interpretable Graph Neural Networks

Abstract

Self-Interpretable Graph Neural Networks have become quite effective at being able to retain the prediction accuracy of comparable black box methods while also introducing interpretability to their decision making processes. However, most self-interpretable methods do not generate stable explanations. In this work, we define a key notion of stability that is absent in the literature and provide a framework that addresses this problem head on. Our framework Information-Theoretic Self-Explainable Graph Neural Network (ITXGNN) is a method that utilizes subgraph extraction with prototypes, clustering, and representation learning in order to satisfy an objective function that prioritizes stability while maintaining accuracy. In addition to our newfound definition of stability we provide theoretical justification of our objective function and derivations highlighting the natural choice of objective function as well as a mathematical bridge from theory to implementation. We conduct comprehensive experiments on several datasets showcasing our method's strength on stability in comparison to leading existing self-interpretable GNNs while maintaining competitive prediction performance. We also conduct ablation studies that break down our frameworks key components and provides quantitative analysis for each components contribution and we also provide a qualitative case study showcasing explanation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.