Linking Representations to Function across Neural Network Populations
Abstract
How do differences in internal representations relate to differences in what intelligent systems can do? The growing diversity of trained neural networks creates an opportunity to address this question by treating representations as a data modality with population-level structure of its own. We introduce a metamodeling framework that learns relationships between representational structure and behavior across systems, using representational measurements as predictive markers analogous to biomarkers in medicine. To compare representations across models with different dimensionalities and coordinate systems, we use symmetry-aware descriptions based on invariant geometric summaries and alignment to shared spaces. Across vision and language, representations elicited by shared stimuli predict variation in visual recognition, robustness to distribution shifts, and linguistic capabilities. These relationships generalize to held-out architecture families in vision and model scales in language, revealing regularities in the mapping from representational organization to function. Our analyses further show that accessing this predictive information depends on how representations are measured: natural images provide more informative probes of visual performance, while linguistic predictions depend on the match between probe stimuli and the capability being assessed. Together, these findings establish internal representations as transferable markers of model capabilities and provide a foundation for a comparative science linking representational structure to function across intelligent systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.