MindLoop: Closed-Loop Brain Supervision for fMRI-to-Image Reconstruction
Abstract
Decoding visual experience from brain activity offers a direct window into how the brain represents the visual world. Modern fMRI-to-image reconstruction decodes cortical activity into conditioning representations for large pretrained generators. However, the targets used to learn these representations are derived primarily from stimulus images and generator-specific objectives: trial-specific brain responses serve as inputs but provide little direct supervision. We refer to this imbalance as supervision asymmetry: it leaves the decoded representation only indirectly constrained by the measured cortical response, allowing a generator to produce plausible images that are weakly grounded in the underlying brain activity and may therefore fail to reflect the individual's own perceptual experience on a given trial. To address this, we introduce MindLoop, which makes the measured cortical response an explicit source of supervision for reconstruction, constraining decoding at three complementary levels: Brain Latent Back-Projection enforces consistency with the same-trial cortical state, Dual-Anchored Relational Alignment preserves stimulus geometry across visual and cortical representations, and Hierarchy-Aware Cortical Fields incorporate retinotopic organization into visual-token readout. On the Natural Scenes Dataset, MindLoop outperforms seven representative methods on seven of eight standard reconstruction metrics, averaged across four subjects. More importantly, brain-response substitution experiments show that reconstruction content tracks the measured cortical response: replacing the brain response with that elicited by a different image collapses identification accuracy to chance and shifts the reconstruction toward the image associated with the substituted response. Cortical fields learned without explicit retinotopic supervision independently recover retinotopic organization, and this organization's horizontal mapping degrades substantially when horizontal-flip augmentation breaks image–cortex correspondence. Together, these results show that balancing supervision from cortical responses and visual stimuli improves reconstruction fidelity, providing a more faithful route toward reconstructing individualized visual experience.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.