Population Similarity Enables Statistically Grounded Comparisons of Representation Alignment
Abstract
Representation similarity metrics are widely used to evaluate neural network alignment. Owing to their quadratic time complexity, these metrics are frequently computed on small samples (), yielding empirical scores that can misrepresent the underlying data-generating process. To capture true alignment, we propose population similarity, i.e., the infinite-sample limit of an empirical score, as the primary target of analysis. We establish the existence of this limit and the consistency of the Triplet and Quadruplet Similarity Indices (TSI and QSI) and Centered Kernel Alignment with an RBF kernel (CKA-RBF), and show that Mutual Nearest Neighbors (MutualNN) admits a population similarity under a mild monotonicity condition. To quantify finite-sample deviations without distribution assumptions, we introduce distribution-free tail bounds, proving non-trivial bounds for TSI, QSI, and CKA-RBF, but not for MutualNN, implying its arbitrarily slow convergence and the failure of empirical scores to reflect its population score. We demonstrate experimentally that the ordering of representation pairs by their finite-sample scores is often inconsistent with their ordering by population similarity, with disagreement rates heavily dependent on the chosen metric and the representation pairs compared. Building on our bounds, we design a one-sided hypothesis test to determine if one population similarity exceeds another by a specified margin, and show its -value can be upper-bounded using only observed scores and sample sizes. These bounds are informative for TSI and QSI at practical sample sizes ( to ). Finally, we showcase the practical utility of our approach with a multimodal case study revealing that CLIP's image and text representations are less aligned at the population level than those of SigLIP2.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.