Sparse Prototype Code Underlies Classification and Prediction Across Modalities
Abstract
Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensional structure makes it difficult to identify the geometric features that determine predictive performance. We identify a universal representational geometry shared across state-of-the-art models in vision, audio, and language, spanning tasks from object recognition to next-token prediction. Its defining feature is that within-class variability is strongly structured: its task-relevant component is aligned with the class’s own centroid and the centroids of competing classes. Building on this observation, we derive an analytical mean-field theory governed by the statistics of projections along true-class and rival-class centroid directions, together with a global renormalization of the class radius that compensates for the non-Gaussianity of real representations. The resulting theory accurately predicts per-class accuracy across tasks, architectures, and modalities. The relevant geometric quantities improve systematically with model scale, revealing how larger models achieve higher accuracy even though their overall within-class variability often increases. A striking feature of this framework is its sparsity: accurate predictions require only a small set of centroid coordinates associated with the true class and its strongest rivals. These rivals are often semantically related to the true class, suggesting a connection to sparse-feature extraction approaches such as sparse autoencoders. Together, these results provide a parsimonious predictive theory of neural representations and suggest that classification and prediction in deep networks are governed by a sparse, centroid-aligned structure embedded within the full high-dimensional representation space.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.