Auditing Explanation Preservation in Sparsified GNN Pipelines
Abstract
We study explanation auditability of graph sparsifiers for graph neural networks (GNNs). Existing sparsifiers preserve structural properties or predictive accuracy, but in deployment workflows where predictions are inspected or audited, do the explanations produced on the sparsified graph remain consistent with the full-graph reference explanations? We introduce an audit framework built around a node-level outcome that permits controlled explanation change and identifies four failure modes: prediction flip, excessive explanation drift, top- structural rearrangement, and top- mass erosion. The framework evaluates preservation at marginal, conditional, and worst-case levels and provides exact finite-sample confidence bounds for the complete test-node population and its structural subgroups. We audit thirteen sparsifiers, four post-hoc explainers, and five node-classification benchmarks spanning homophilic and heterophilic graphs. Experiments spanning 280 main audit runs indicate that (i) prediction-preserving sparsification does not imply explanation-preserving sparsification, (ii) distinct sparsifier families exhibit signature failure modes, (iii) preservation rates vary systematically with local graph structure, with high-degree nodes substantially less likely to satisfy the audit outcome, and (iv) overall preservation rates do not establish that the same preservation target holds across all structural groups. We argue that explanation preservation is a measurable property of deployed GNN pipelines, distinct from accuracy, and one that should be reported alongside it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.