acceptodds
Under review as a conference paper at ICLR 2027

Modeling Local, Global, and Cross-Modal Context in Multimodal 3D MRI

Abstract

Brain MRIs pose a fundamental challenge for machine learning: models must learn from high-dimensional 3D data spanning multiple co-registered modalities, despite the limited sample sizes typical of neuroimaging studies relative to the diversity in anatomy, pathology, and acquisition conditions. While multimodal imaging provides complementary information critical for clinical interpretation, effectively integrating these signals remains difficult. We propose Multimodal Intra- and Cross-Context Vision Transformer (MICViT), a 3D Vision Transformer that explicitly models both modality-specific representations and cross-modal interactions across local and global contexts. Concretely, MICViT combines four attention mechanisms: modality-specific local and global attention for intra-modal feature learning, and cross-modal local and global attention to capture interactions between modalities. We evaluate MICViT on brain age prediction across four heterogeneous datasets (UK Biobank, SOOP, ADNI, Cam-CAN) using multiple MRI modalities (e.g., T1, FLAIR, DWI, SWI). MICViT consistently achieves better results than established CNN and Transformer baselines in 3D settings. Notably, it benefits more strongly from multimodal inputs, yielding larger performance gains as additional modalities are incorporated. Beyond age prediction, we further evaluate the learned representations on disease-onset prediction in UK Biobank and Alzheimer’s disease classification in ADNI, where MICViT shows competitive to strong performance across clinically relevant endpoints. These results suggest that explicitly modeling intra- and cross-modal interactions is a promising direction for representation learning in multimodal neuroimaging. Code is available at https://anonymous.4open.science/r/MICViT-C057

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.