LAVA: Explainability for Unsupervised Latent Embeddings
Abstract
Unsupervised black-box models are drivers of scientific discovery, but are difficult to interpret as their output is often a multidimensional embedding rather than a well-defined target. While explainability for supervised learning estimates how input determines model output, its unsupervised counterpart should uncover how input is captured by the learned latent structure. However, existing adaptations of supervised model explainability for unsupervised learning deliver explanations either at the single-sample or the dataset-summary level, too fine-grained or reductive to be meaningful, and cannot explain models without a mapping function. To address this, we propose LAVA, a model-agnostic method to explain embedding structure through feature covariation in the input data. Using prototype-like explanations, LAVA captures local patterns of input feature correlations that may reoccur globally across the embeddings. LAVA offers stable explanations at a chosen granularity, revealing domain-relevant patterns such as pixel motifs in images and disease signals in cellular processes otherwise missed by existing methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.