MindStrata: Learning Robust Multi-Subject Cortical Visual Representations via Stratified Layout-Semantics Decoding
Abstract
Decoding cortical visual representations that generalize across human subjects remains challenging due to (i) limited modeling of compositional semantics and structured layout in multi-object scenes, (ii) brain functional differences across subjects, and (iii) evaluation protocols dominated by feature-based scores that are often misaligned with human perception. To address these challenges, we propose __MindStrata__, a neuro-cognitively inspired framework for multi-subject brain-to-scene decoding via stratified modeling. __MindStrata__ maps fMRI into structured brain-scene representations using two complementary branches: a ventral-like branch decodes multi-level semantics (objectinstancesscene) in a CLIP-aligned space, while a dorsal-like branch decodes multi-scale spatial organization (targetsurroundingglobal) within latent layout space. A Brain-Scene Routing Controller integrates these stratified semantic and spatial cues for structure-guided reconstruction. To mitigate multi-subject variability, we further introduce Consistency-based Functional Alignment, combining subject-specific normalization with a relational consistency loss. We also propose the __Semantic Faithfulness Evaluation Protocol__ that evaluates description-, category-, count-, and position-level correctness and aligns more closely with human preferences Experiments demonstrate fine-grained layout and semantic decoding across single-subject, multi-subject, and data-limited settings, reveal ROI-specific spatial localization, and show robust reconstruction performance under progressively reduced fMRI spatial resolution. Code will be publicly released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.