Towards Neurosymbolic Mechanistic Interpretability: The Case of Graph Neural Networks
Abstract
Mechanistic interpretability (MI) aims to reverse-engineer subgraphs of neural networks into human-understandable circuits; yet, it lacks a formal language for specification and verification. This often results in leaky abstractions and unresolved theoretical problems with feature identifiability and universality. While symbolic AI generally avoids such pitfalls by leveraging the language of formal logic, it is naturally limited in both scope and scale. We propose a neurosymbolic synthesis, using MI to decompile latent representational superposition into discrete symbolic primitives, shifting interpretability from heuristic trial-and-error to a rigorous, logic-guided process. We demonstrate the feasibility of this synthesis through a case study on graph neural networks, utilizing their known logical expressiveness bounds to mathematically constrain the vocabulary of the respective latent feature space. Our experiments show that MI techniques, particularly sparse autoencoders, can reliably extract relational predicates, with targeted interventions verifying their logical compositions and interpretations. Ultimately, this work opens a path toward formal symbolic discovery in (relational) neural models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.