Procrustes Similarity with the Reconstructed Space: A Framework for Representational Similarity and Relative Information Content
Abstract
With the growing reliance on high-dimensional embeddings in many areas including machine learning, neuroscience, and computational biology, the need for rigorous methods to compare and assess these representations and their information content is critical. Quantifying the similarity between two variables or distributions is a longstanding challenge in statistics and information theory, and this challenge becomes even more pronounced in high-dimensional spaces. Established methods, such as Mutual Information (MI), Canonical Correlation Analysis (CCA), and Centered Kernel Alignment (CKA), struggle with accurately capturing complex dependencies, handling noise, and modeling nonlinear relationships in high-dimensional settings, limiting their practical usability. To address this, we introduce Procrustes Similarity with the Reconstructed Space (PSRS), a novel approach that quantifies the Relative Information Content (RIC) between two embedding spaces by reconstructing one from another using deep learning techniques and measuring the reconstruction fidelity via Procrustes similarity. Through a combination of synthetic tests and real-world case studies, we demonstrate that PSRS significantly outperforms state-of-the-art methodologies in capturing both linear and nonlinear relationships while exhibiting superior robustness to noise and extraneous information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.