CortexLink: Reconstructing Long Videos from fMRI with Event-Level Semantic Consensus and Continuous Geometry
Abstract
Current fMRI-to-video systems typically reconstruct continuous visual experience as independent short clips, overlooking the event structure shared by neighboring neural measurements. This mismatch discards repeated semantic evidence and amplifies local uncertainty into content drift and geometric discontinuities across clip boundaries. We introduce CortexLink, an event-aligned framework that reorganizes clip-level predictions within stimulus shots. Its semantic branch treats decoded descriptions as noisy observations of a shared event and distills them into a consensus anchor through Minimum Bayes Risk decoding. Its geometric branch combines shot-level spatial aggregation during training with groupwise shape–pose decomposition at inference, recovering a shared subject template and a continuous pose trajectory. The resulting semantic and geometric anchors jointly guide long-video synthesis and reset at event boundaries. Experiments show consistent improvements in semantic fidelity, motion agreement, and cross-clip continuity over clip-level baselines. On shots lasting at least 10 seconds, CortexLink improves 50-way identification by \(15.5%\), while reducing Dynamic Degree error by \(51.8%\) and Boundary Seam by \(67.3%\) relative to CortexVideo. Component analyses further verify the complementary contributions of semantic consensus and continuous geometry. These results demonstrate that event structure provides an effective organizing principle for advancing fMRI reconstruction from isolated clips toward coherent long videos.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.