Quantifying the Accuracy-Interpretability Trade-off in Concept Sidechannel Models
Abstract
Concept bottleneck models make predictions through human-understandable concepts, providing interpretability. However, this bottleneck limits task accuracy when the concept set is incomplete or the concept-to-task map is insufficiently expressive. Recent concept-based models address this limitation by adding sidechannels that transmit additional information around the bottleneck. These sidechannels often recover accuracy, but they also raise a basic diagnostic question: how much can the final prediction still be interpreted using the concept pathway, and how much interpretability is lost? We study this question through a unified probabilistic view of concept sidechannel models. The framework exposes two inference modes: the default inference mode of the model, and a bottleneck mode, in which the sidechannel is replaced by an input-independent distribution. Comparing these modes yields an architecture-agnostic invariance metric, and naturally leads to a regularizer that penalizes discrepancies between the two modes. Across several architectures and datasets, we find that accuracy-only training often produces models whose predictions cannot be reproduced without the sidechannel, even when concept-only prediction is feasible. SIS regularization allows users to trace a controlled accuracy-interpretability frontier, improves responsiveness to concept interventions, and can yield task predictors that rely more visibly on concepts. Our results provide practical tools for auditing and tuning the role of sidechannels in concept-based models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.