acceptodds
Under review as a conference paper at ICLR 2027

SIREM: Speech-Informed MRI Reconstruction with Learned Sampling

Abstract

Reconstructing speech rtMRI from highly undersampled acquisitions remains challenging because the rapid articulatory dynamics of speech demand both high spatial and temporal resolution. Yet speech rtMRI is acquired together with an information source that conventional reconstruction methods largely ignore: the synchronized acoustic signal. We introduce SIREM, a speech-informed MRI reconstruction framework that exploits this cross-modal relationship as an auxiliary prior. Because vocal-tract configurations are closely coupled to the acoustics they produce, part of the image content can be inferred from speech while the acquired k-space data provide complementary anatomical information. SIREM therefore decomposes each reconstructed frame into audio-driven and MRI-driven components that are combined through an anatomically derived spatial weighting map. The audio branch predicts articulator-related structure from synchronized speech, whereas the MRI branch recovers complementary content from the measured k-space samples. We further introduce a learnable soft weighting profile over spiral arms, allowing the contribution of individual k-space arms to be optimized jointly with the multimodal reconstruction. This provides a unified formulation for audio-guided prediction, MRI reconstruction, and k-space arm reweighting. We evaluate SIREM on the USC 75-Speaker speech rtMRI benchmark against conventional reconstruction approaches, including wavelet-based compressed sensing, total variation, and a deep learning method. SIREM achieves 44–53 higher reconstruction throughput than the iterative methods while preserving anatomically plausible vocal-tract structure, and achieves real-time computational throughput, with per-frame compute within a 30-fps budget. Our results provide an initial benchmark for speech-informed rtMRI reconstruction and demonstrate the potential of synchronized acoustics as an auxiliary prior for accelerating dynamic speech MRI.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.