Measuring Model–Brain Alignment in Latent Spaces
Abstract
Evaluating neural network models as hypotheses of brain computation requires accurately quantifying the alignment between artificial and biological representations. Here, we propose measuring representational alignment within a shared latent space learned via a cross-modal multi-encoder–decoder architecture. By mapping network and brain representations into a common low-dimensional space, our framework enables direct comparison of both overall latent distributions and per-stimulus correspondences. To demonstrate the properties of our method, we construct two diagnostic benchmarks using fMRI data paired with reference and manipulated network representations: (1) structural "foil" models, which are mechanistically dissimilar to brain computation by design yet match its coarse high-level structure, and (2) reference model representations with systematically corrupted information. We show that linear ridge regression is nearly flat across these structural transformations, whereas commonly used similarity metrics – Representational Similarity Analysis (RSA) and Centered Kernel Alignment (CKA) – score the foil models close to, or above, the reference network. In contrast, metrics derived from our latent space (i) accurately capture progressive information loss under controlled corruption and (ii) reliably identify structural discrepancies between the foil models' representations and biological neural data, providing a new robust, structure-sensitive tool for brain–AI alignment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.