acceptodds
Under review as a conference paper at ICLR 2027

Towards Neurosymbolic Mechanistic Interpretability: The Case of Graph Neural Networks

Abstract

Mechanistic interpretability (MI) aims to reverse-engineer subgraphs of neural networks into human-understandable circuits; yet, it lacks a formal language for specification and verification. This often results in leaky abstractions and unresolved theoretical problems with feature identifiability and universality. While symbolic AI generally avoids such pitfalls by leveraging the language of formal logic, it is naturally limited in both scope and scale. We propose a neurosymbolic synthesis, using MI to decompile latent representational superposition into discrete symbolic primitives, shifting interpretability from heuristic trial-and-error to a rigorous, logic-guided process. We demonstrate the feasibility of this synthesis through a case study on graph neural networks, utilizing their known logical expressiveness bounds to mathematically constrain the vocabulary of the respective latent feature space. Our experiments show that MI techniques, particularly sparse autoencoders, can reliably extract relational predicates, with targeted interventions verifying their logical compositions and interpretations. Ultimately, this work opens a path toward formal symbolic discovery in (relational) neural models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.