Archetypes of a Concept: Sample-Traceable Decomposition of Concept Representations
Abstract
Concept-direction methods summarize a concept with a single vector. This representation is compact, but it obscures variation within the concept and the training examples that support that variation. We introduce Archetypes of a Concept (AoC), a post-hoc method that decomposes concept-positive features into simplex-constrained archetypes. Each archetype is a convex combination of training features and therefore retains explicit sample-level provenance. A fixed concept probe then selects aligned archetypes. New queries are represented by simplex coefficients over the selected dictionary. We establish convex-hull traceability and a margin-based condition for preserving fixed-probe predictions. We also analyze coefficient continuity and selection stability under explicit regularity and score-gap conditions. Together, these properties support evidence inspection and coefficient-space analysis in a frozen feature space.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.