CABM: Efficient Concept Augmentation for Interpretable Deep Models
Abstract
Concept Bottleneck Models (CBMs) offer interpretability by aligning latent representations with semantic concepts. However, standard CBMs suffer from a static concept space, limiting their adaptability to evolving knowledge. Expanding this space typically necessitates full retraining, which is computationally prohibitive and risks semantic drift in pre-existing concepts. In this work, we propose Concept Augmentation in Bottleneck Models (CABM), a novel framework for post-hoc concept addition that avoids iterative end-to-end retraining. CABM formulates concept addition as a parameter perturbation problem, addressed by influence-inspired local updates. To overcome the scalability bottleneck of dense curvature inversion in high-dimensional spaces, we leverage Eigenvalue-corrected Kronecker-Factored Approximate Curvature (EK-FAC) to perform structured approximate second-order updates. We further introduce a trust-region constraint and concept-conditional initialization to improve numerical stability. We bound its local error relative to a re-optimized appended concept block, including the effect of trust-region projection. Extensive experiments show that our framework achieves predictive performance on par with full retraining while reducing computational costs by – relative to full retraining and – relative to a matched converged frozen-feature probe, offering a robust solution for dynamic interpretable modeling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.