Are Dictionary Learning Metrics for Neural Concepts Interchangeable Proxies for Human-Interpretability?
Abstract
Concept-based methods aim to make neural networks more interpretable by representing their behavior in terms of human-meaningful concepts. As learned concept representations have become more common, recent work has proposed reconstruction fidelity, sparsity, Representation Diversity and Stability metrics to evaluate their quality. However, it remains unclear whether such metrics are actually informative of how comprehensible learned concept spaces are to humans, a prerequisite for increasing neural network interpretability. We study this question empirically by adopting an activation-space factorization view of learned concept representations, following recent work that frames concept extraction as dictionary learning. We then test whether dictionary learning metrics of the resulting concept spaces are associated with human evaluation of those spaces through a study with four complementary tasks and five measures of concept comprehension, involving 351 participants and yielding 5,616 individual responses. Our results show that different metrics exhibit substantially different relationships to human evaluation across tasks, and that these relationships are robust to the influence of individual classes. While reconstruction-based and stability metrics show no consistent connection to human judgments, Gini Sparsity and Representation Diversity display associations with linguistic comprehension measures for the studied setting. Overall, standard dictionary-learning metrics do not provide sufficient evidence for the evaluation of human-comprehensible concept spaces and should not be treated as interchangeable proxies for interpretability. However, some of them can be used as candidate indicators thereof in a systematic, human-validated analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.