acceptodds
Under review as a conference paper at ICLR 2027

FisherCBM: Concept Bottlenecks with Closed-Form Concept Transport and Fisher-Bounded Residual

Abstract

Concept bottleneck models route a prediction through a layer of annotated concepts, but in practice the layer transmits information the concepts do not encode, and nothing determines which part of the input a concept is read from. Both problems arise because each concept is predicted from the entire input. We present FisherCBM, in which each concept is computed from an explicit set of input tokens. A concept transports the token features of the input, such as the patch tokens of a Vision Transformer, toward each of two learnable prototypes through an unbalanced optimal transport problem with a closed-form solution: the plan toward each state is the concept's saliency map and its mass is the evidence for each state. The head receives a one-bit state and a residual gated by an evidential belief and bounded by a fixed constant. The residual's sensitivity to a token is proportional to that token's share of the plan, its token Jacobian is explicit, and the resulting Fisher penalty on off-support tokens costs a forward pass. The bounded residual gives every intervention a guaranteed direction and a lower bound on its magnitude, and reduces whether the prediction follows to a sign condition on the head. On CUB and AwA2, FisherCBM matches CEM in task accuracy with one fifth of its trainable parameters and rises monotonically to 100% accuracy under correct interventions, while wrong interventions drive it to chance as they do a hard CBM; it attains the highest concept accuracy on CUB and CelebA, and on CelebA, where two label attributes are not concepts, the bounded residual passes only part of the missing information.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.