acceptodds
Under review as a conference paper at ICLR 2027

Beyond Concept Generation: Interpretable Concept Scoring for Debiasing CLIP-based CBMs

Abstract

Concept Bottleneck Models (CBMs) explain predictions through human-readable concepts. However, label-free CBMs built on CLIP image-text similarities may inherit biases from CLIP representations, allowing visually or contextually associated concepts to serve as semantically unreliable shortcuts. Addressing such shortcuts is non-trivial: a concept that provides valid evidence for one class may form a misleading association with another. Therefore, globally removing potentially problematic concepts can be harmful, as it may discard useful information for other classes. To address this challenge, we introduce ScoreCBM, an interpretable concept-scoring framework for debiasing CLIP-based CBMs. Moving large language models (LLMs) beyond their conventional role of concept generation, we use them to explicitly score the semantic relevance of class-concept pairs. The resulting scores provide a class-specific semantic prior that guides two complementary interventions: score-guided activation attenuation suppresses low-relevance concept activations during projection learning, while adaptive regularization penalizes low-relevance concept-class connections in the final sparse classifier. Our framework requires neither group labels nor modifications to CLIP, while providing a traceable and editable interface for model intervention. To scale to large datasets, we further introduce CLIP-based top-\(r\) pre-filtering, reducing LLM scoring complexity from quadratic to linear in the number of classes. Experiments across biased and standard benchmarks demonstrate consistent improvements in predictive performance. On BAR, our method improves LF-CBM accuracy from 84.25% to 89.39%, surpassing the 88.91% non-interpretable ERM baseline. The framework further scales to the 1,000-class ImageNet, improving LF-CBM from 71.87% to 75.28%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.