acceptodds
Under review as a conference paper at ICLR 2027

Archetypes of a Concept: Sample-Traceable Decomposition of Concept Representations

Abstract

Concept-direction methods summarize a concept with a single vector. This representation is compact, but it obscures variation within the concept and the training examples that support that variation. We introduce Archetypes of a Concept (AoC), a post-hoc method that decomposes concept-positive features into simplex-constrained archetypes. Each archetype is a convex combination of training features and therefore retains explicit sample-level provenance. A fixed concept probe then selects aligned archetypes. New queries are represented by simplex coefficients over the selected dictionary. We establish convex-hull traceability and a margin-based condition for preserving fixed-probe predictions. We also analyze coefficient continuity and selection stability under explicit regularity and score-gap conditions. Together, these properties support evidence inspection and coefficient-space analysis in a frozen feature space.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.