acceptodds
Under review as a conference paper at ICLR 2027

Reliable Neuron Identification: A Statistical Learning Framework for Neuron Explanations

Abstract

A central goal of mechanistic interpretability is to reverse-engineer deep neural networks into human-understandable mechanisms. A foundational step in this pipeline is neuron identification, which maps internal neural representations to human-understandable concepts. However, existing methods for neuron identification largely operate as empirical heuristics, making it unclear when the resulting neuron explanations can be trusted. To bridge this critical gap, we establish the first theoretical foundation for neuron identification by uncovering a fundamental correspondence to statistical learning theory: for a given neuron, selecting the concept that best explains its activation behavior is equivalent to Empirical Risk Minimization (ERM) over a concept hypothesis class. This ERM formulation unifies popular identification methods (e.g., Network Dissection, CLIP-Dissect) under a common objective and casts explanation reliability as a classical generalization problem. Leveraging this connection, we derive PAC-style fidelity bounds that characterize how explanation reliability depends on probing sample size and concept class complexity, and reveal a critical limitation of commonly used metrics: AUROC- and recall-based identification can become substantially less reliable for rare concepts. Finally, to capture uncertainty in neuron explanations, we introduce Bootstrap Explanation (BE), a principled procedure that produces concept prediction sets, with a provable coverage guarantee for its idealized variant. By connecting interpretability with statistical learning theory, our work provides a principled foundation for reliable neuron identification.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.