CircuitFACT: Circuit Discovery for Mechanistic Interpretability of Graph Neural Networks
Abstract
Mechanistic interpretability seeks to understand neural networks by identifying the internal computational mechanisms underlying their behavior, often represented as circuits. However, recent circuit discovery methods have focused predominantly on transformer language models, and their application to graph neural networks (GNNs) remains largely unexplored. We propose CircuitFACT, a circuit discovery method for GNNs built on computation trees, which explicitly represent the message-passing dependencies involved in a GNN prediction. CircuitFACT first constructs these computation trees, attributes their components using Shapley-based attribution, and trims the resulting structures to obtain compact circuits. These circuits provide a structured view of the computations supporting a prediction and can be related back to the corresponding elements of the input graph.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.