Gene2Brain-AE: Self-Supervised Representation Learning for Genetically-Informed Patterns of Brain Atrophy
Abstract
Characterizing heterogeneity in neurodegenerative diseases from imaging alone is limited, as similar patterns of brain change can arise from distinct underlying mechanisms. Integrating genetic information offers a way to disentangle these processes, but genetic effects are indirect, partially mediated, and entangled with non-genetic sources of variation. We propose Gene2Brain-AE (G2B-AE), a self-supervised multimodal framework that combines an interpretable NMF-style imaging autoencoder with a genetic variational autoencoder, coupled through a directional predictor that maps genetic embeddings into a subset of imaging latent factors. This directional, predictive coupling reflects the asymmetry of the underlying biological pathway and decomposes imaging variation into genetically predictable and residual components, without requiring genetic input at inference time. On semi-synthetic data, G2B-AE recovers genetic and residual patterns that symmetric alignment-based variants entangle or absorb into the wrong subspace. Applied to ADNI, the model identifies multiple genetically-guided atrophy axes, including distinct APOE-associated patterns, alongside a residual axis linked to vascular and metabolic variables; despite using no genetic input at test time, these components' clinical, cognitive, and neuropathological signatures replicate across eight independent external cohorts, indicating generalizable biology rather than dataset-specific correlations. These results show that directional, partially shared latent representations offer a more biologically-informed and interpretable characterization of disease heterogeneity. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.