Angular Density Represents an Embedding's Typicality
Abstract
Self-supervised pretraining has become the dominant paradigm for learning general-purpose representations across image, multi-modal, and language settings. Whether a given input is well-represented is normally assessed only indirectly, through labels, class structure, or expensive downstream evaluation. We show that an embedding's own geometry offers a cheap, label-free substitute: an embedding's *angular density* and, less consistently, its norm predict representation quality across paradigms. We argue that angular density measures an embedding's local *typicality* and does so even where no label- or class-based notion of typicality is available to begin with, e.g., individual LLM prompts or one side of a cross-modal image-text pair. Across vision, vision-language, and language models, embeddings with high angular density are more likely to be classified or answered correctly and, in vision and vision-language settings, represented similarly by other models. We additionally find that, in multi-modal models, the geometry of a *text* embedding predicts how reliably the corresponding *image* embedding is classified, despite no parameters being shared between the two encoders. We show that this local signal succeeds where a natural alternative, distance to a dataset centroid, does not: the global notion of typicality is inconsistent across models and often inverts, while angular density is monotonically predictive throughout.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.