The complexity of understanding: Formalizing Concept Discovery in Mechanistic Interpretability
Abstract
Mechanistic Interpretability (MI) is an emerging subfield of eXplainable Artificial Intelligence (XAI), which aims to reverse-engineer the internal representations of black-box machine-learning models (like neural networks) into human-understandable concepts. Despite increasing attention and promising results, the literature still lacks a formalization of the most basic and fundamental MI concept of "feature", leading to limited rigor and inconsistencies in the general adoption of MI, as well as a continued reliance on human judgment for its evaluation. In this work, we fill this gap by devising, for the first time, a formal definition of the notion of mechanistic feature, grounded in the core principles of the MI literature. We formulate a novel optimization problem to soundly identify mechanistic features according to our definition, and theoretically characterize it in terms of NP-completeness, inapproximability, and mixed-integer nonlinear programming (MINLP) formulation. To face the inherent computational complexity of the problem, we design two well-founded heuristics: one consisting in solving a relaxed yet easier- and faster-to-solve MINLP formulation of our problem, and the other relying on a combined local- and beam-search strategy. We extensively evaluate our methods on numerous datasets and testing scenarios. The obtained results demonstrate their high effectiveness and superiority over a variety of baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.