acceptodds
Under review as a conference paper at ICLR 2027

When are Concept Bottleneck Model Explanations Faithful and Compact?

Abstract

Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability – we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants, faithful explanations must include all concepts in the bottleneck, compromising interpretability when this is large. This result applies to both heuristic and faithful-by-construction formal explanations. To ensure compact faithful explanations exist, we suggest i) modeling concepts probabilistically as binary or categorical random variables (rather than logits), and ii) employing per-concept training-time sparsification via group lasso (rather than regular elastic net). We also extend algorithms from formal explainability to CBMs, and show they outperform heuristics in terms of guarantees and explanation size. Overall, our work warns against naıve interpretability claims and provides formal conditions and practical strategies for encouraging CBMs to be as interpretable as advertised.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.