Hierarchical Concept Bottleneck Large Language Models for Controllable Generation
Abstract
Concept Bottleneck Models (CBMs) provide an interpretable interface for inspecting and intervening on model predictions through human-understandable concepts. For language generation, a key challenge is to support control at different levels of semantic granularity while ensuring that interventions on concept representations effectively influence generated text. To address this challenge, we introduce Hierarchical Concept Bottleneck LLMs (H-CBLM), a framework for controlling language generation through coarse-grained and fine-grained concepts. H-CBLM uses coarse-grained concept logits to gate fine-grained concept logits according to a predefined hierarchy, while retaining residual features to support language modeling. We further introduce an intervention reconstruction objective that updates only the concept-to-token weights to reconstruct text after label-based concept replacement, complementing residual regularization adapted to partially observed multi-label supervision. At inference time, users can intervene on fine-grained concepts representing specific attributes or coarse-grained concepts representing broader semantic categories. Compared with the adapted CB-LLM baseline, H-CBLM increases the success rate of generation satisfying both target-concept and quality criteria by an average of 32.74 percentage points across six datasets, while coarse-grained interventions yield more evenly distributed expression of associated fine-grained concepts across all four evaluated datasets. In an LLM agent application, local coarse-grained interventions reduce tool-call attempts by 46.96%, at the cost of a minor 2.33-percentage-point decrease in task accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.