acceptodds
Under review as a conference paper at ICLR 2027

Matryoshka Concept Bottleneck Models

Abstract

Concept bottleneck models predict through named attributes that can be inspected and replaced. Several readouts can give a frozen encoder different input budgets. We study joint training that makes those budgets part of learning the encoder itself. Matryoshka Concept Bottleneck Models (MCBMs) optimize one concept encoder with task losses on nested prefixes of a fixed concept order. Each prefix has a trained head and a stable set of concept identities. At deployment, a query policy operates within the selected prefix. On CUB, the same model reaches accuracy with 64 concepts and with all 112, reducing visible inputs by . Single-seed studies on CelebA and the animal domain of LAD show similar retention at moderate budgets. Readout studies on the same CUB encoders examine a separate property: response to concept replacements. Active queries improve probability-head curves, while intervention-aware training improves raw-head responses. These results support joint multi-budget training as a practical design for named concept interfaces. They also separate task accuracy at reduced budgets from the ability of a particular head to use replacement values.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.