Localizing Memorization in Graph Neural Networks
Abstract
Graph Neural Networks (GNNs) have been shown to memorize individual training examples, with recent work quantifying their memorization. However, where memorization occurs within the network, i.e., which layers and weights are responsible, remains an open question. While studies in Deep Neural Networks (DNNs) have established that memorization often concentrates on a small number of model weights across all model layers, no such analysis has been explored for GNNs. This paper presents the first systematic study of localizing memorization in GNNs at both the layer and weight levels. At the layer level, we employ gradient accounting to measure per-layer gradient norm contributions from memorized and non-memorized nodes. At the weight level, we analyze memorization of a training node by measuring the minimum number of important weights whose ablation leads to the node's prediction flip. By localizing memorization in multiple GNN architectures trained on diverse datasets of varying homophily levels, we find that (1) memorized nodes dominate the gradient contributions over non-memorized nodes at every layer, with the gap generally increasing with layer depth, (2) the graph homophily determines whether memorized nodes produce gradients aligned with or opposed to those of non-memorized nodes, revealing a GNN-specific memorization mechanism, (3) the memorization of individual nodes in GNNs is concentrated on a few model weights, and (4) finally, the memorization localization analysis in GNNs enables targeted mitigation of privacy leakage. Overall, our work contributes to the fundamental understanding of memorization in GNNs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.