acceptodds
Under review as a conference paper at ICLR 2027

BrainWM: A Multimodal World Model for Representing, Forecasting, and Simulating Whole-Brain Dynamics

Abstract

Most brain imaging foundation models treat entire fMRI scans as static images, yet the brain is a dynamic world: its state is described by both structure and function, and it evolves, driven by intrinsic spontaneous activity, external stimuli, and long-term processes such as aging. We present BrainWM, a multimodal world model for whole-brain 4D dynamics that unifies representation learning, forecasting, and simulation in a single model with three components: BrainWM-VAE encodes structural and functional MRI into an integrated structure–function latent state; BrainWM-DiT learns to generate the next state from the past in this latent space via flow matching; and BrainWM-BAM learns latent actions for state transitions through prior–posterior modeling. BrainWM is trained progressively in three stages: first unified latent-space learning, then within-run generative forecasting pretraining, and finally post-training for conditional cross-run simulation. On large-scale internal and external test sets, BrainWM achieves performance on downstream tasks that surpasses or is comparable to that of foundation models designed specifically for representation learning, while also being capable of forecasting and simulating functional and structural brain changes—tasks that these models are not designed to address. To our knowledge, BrainWM is the first brain imaging world model systematically evaluated for representing, forecasting, and simulating whole-brain dynamics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.