Know What to Trust and What to Check: Calibration-Free Auditing for Concept Bottleneck Models
Abstract
Concept bottleneck models make predictions through interpretable intermediate concepts. When their concept predictions are uncertain, a reviewer needs to know whether this uncertainty can change the decision and which concept to inspect. We study these questions using an input-dependent ellipsoid constructed from ensemble disagreement. For a linear task head, a closed-form margin certifies that the predicted class is uniquely optimal throughout the ellipsoid and its intersection with the probability box. For diagonal geometry, a second diagnostic decomposes the squared first-order sensitivity penalty into concept contributions. Its Shapley interpretation concerns a fixed local surrogate, rather than the task loss after an observed correction. We evaluate the margin and practical ranking rules on concept prediction tasks, separating them from the optional robust training procedure. The reported margin-based selective prediction scores are comparable to confidence and entropy, while training results vary across datasets. On CEBaB, single-concept oracle replacement lowers accuracy under every ranking rule evaluated, including random selection. Additional five-seed CEBaB ablations with matched dropout and initialization give similar accuracy, ambiguity width, and margin-based retained accuracy. The construction provides a conditional account of model sensitivity; it does not establish coverage of the true concepts or guarantee that an intervention will correct a
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.