NAMED FROM THE START: CONCEPT-ANCHORED SPARSE AUTOENCODERS FOR INTERPRETABILITY OF VISION-LANGUAGE MODELS
Abstract
Sparse autoencoders (SAEs) decompose model activations into interpretable features, but learn those features without semantic identities. Names are assigned only afterward, through inspection or similarity matching, and this step often fails even when the desired concept is present in the learned dictionary. We hypothesize that the problem lies not in feature discovery but in semantic assignment: post-hoc methods must align an already learned dictionary with language, although nothing during training required their geometries to agree. We introduce NAMED-SAE, a concept-anchored SAE that makes semantic identity part of dictionary learning. Text-derived concept embeddings define the named latents, a shared orthogonal adapter aligns them with the activation space, and bounded concept-specific residuals permit adaptation to visual evidence without unrestricted semantic drift. A separately budgeted bank of unnamed latents captures information outside the vocabulary. The adapter is estimated once and frozen; it can be fitted using limited labelled examples or inferred without annotations from unsupervised pseudo-means. Across CLIP and SigLIP, \name provides complete coverage of the target concepts, produces names that are more faithful to activating images and more predictive of intervention effects than post-hoc assignments, and retains most of the reconstruction quality of an unconstrained SAE. These identities also persist under distribution shift without retraining or renaming, while changes in latent activity describe the semantic content of the shift itself. Our results show that semantic identity is more reliably imposed as a constraint during feature learning than recovered through analysis afterward.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.