Conceptual Archetype Decomposition for Tracing Reconstructed Neural Activations
Abstract
Concept decompositions of neural activations usually expose a dictionary and reconstruction weights, but inspecting how a reconstructed test feature draws on particular training locations requires an explicit provenance map. We present Conceptual Archetype Decomposition (CAD), a post-hoc explanatory interface derived from classical archetypal analysis. Concepts are convex combinations of full-image training activations, and test features are reconstructed as convex combinations of these fixed concepts. Composing the two coefficient matrices gives a normalized map from each test location through concepts to training locations. This map is exact for the learned decomposition, remains available before optimizer convergence, and its reconstructed features satisfy a perturbation bound when the coefficients are fixed. Existing controlled CUB results describe reconstruction sensitivity to the implementation choices; we distinguish these results from a secondary pairwise logit diagnostic that does not establish prediction fidelity. CAD supplies inspectable training evidence while leaving the original predictor unchanged. The core implementation code is available in the supplementary material.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.