On the Fragility of Representational Similarity Metrics
Abstract
Representational similarity metrics are widely used tools that measure correspondences between machine learning (ML) models' learned representations. We show that these metrics can be brittle and raise important questions about whether they provide meaningful notions of model similarity. To do this, we introduce an adversarial optimization framework that manipulates the parameters and thus the representation space of an adversarial target model relative to a reference benign model. Using this framework, we show that we can inflate or suppress representational similarity metrics, such as Centered Kernel Alignment (CKA) and Singular Vector Canonical Correlation Analysis (SVCCA), in the adversarial model while keeping its parameters and task performance close to those of a benign model. We can also make the adversarial and benign models functionally unaligned but keep their representational similarity high and parameters similar. Across five datasets and eight architectures, our findings show that high functional similarity between models does not necessarily imply high representational similarity between them, and vice versa.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.