Concept Envelope Model
Abstract
Pretrained embeddings can be highly predictive while encoding task-relevant concepts in a form that is not readable as a single direction. Yet concept-bottleneck and dictionary-score methods often expose each named concept through a scalar score, which can obscure semantic structure spread across multiple embedding directions. Motivated by the Linear Representation Hypothesis and semantic structure in embedding space, we introduce Concept Envelope, which connects predictive sufficiency with concept-based interpretability and offers a way to retain task-related semantic variation beyond prediction alone. Concept Envelope is a supervised neural method that jointly learns a concept subspace and a nonlinear predictor with soft covariance separation, followed by post-hoc semantic grounding. Across seven visual benchmarks, it maintains competitive predictive accuracy. Two quantitative interpretability metrics show retention of independently defined semantic directions, including those missed by accurate prediction-only subspaces, and transfer of estimated edits to held-out intervention pairs. Controlled language tasks further show that dictionary scores can follow surface wording rather than latent state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.