MindRefine: Language-Informed Semantics and Perceptual Refinement for Cross-Subject Brain Visual Decoding
Abstract
Cross-subject functional magnetic resonance imaging (fMRI) visual decoding remains challenging when only limited paired data are available for a target subject. Existing methods mainly align brain responses with pretrained visual representations, leaving semantic structure only implicitly constrained, and rely on reconstruction pipelines that lack sample-specific perceptual correction. We propose MindRefine, a framework that improves limited-data fMRI visual decoding through representation learning and reconstruction refinement. Language-informed semantic regularization uses structured descriptions generated by a vision-language model (VLM) exclusively from training images to provide gated multi-channel semantic supervision and semantic-neighborhood distillation. This introduces explicit semantic constraints without using language information at inference. Feature-space reconstruction refinement keeps the trained decoder frozen, estimates sample-specific perceptual and semantic targets from a paired training reference bank, and refines each reconstruction through stage-wise feature-space optimization. Extensive experiments on the Natural Scenes Dataset in the limited-data regime demonstrate that MindRefine improves both low-level reconstruction fidelity and high-level perceptual consistency while preserving competitive retrieval performance. Additional experiments under the full-data setting further confirm the effectiveness and robustness of the framework. The code is available at https://anonymous.4open.science/r/MindRefine.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.