SFG-Mamba: Reliability-Aware Speech Frame Graph Propagation with State-Space Modeling for Speech Enhancement
Abstract
Single-channel speech enhancement (SE) remains challenging when degradation exhibits temporal heterogeneity. Most existing methods process speech as temporally ordered sequences, without frame-level reliability estimation or direct reuse of information from reliable frames across time. To address this problem, we propose SFG-Mamba, a framework that incorporates a reliability-aware Speech Frame Graph (SFG) to model reliability-guided cross-temporal information transfer. A Speech Time-Frequency Graph Convolution (STGC) module further integrates non-local graph information with local time-frequency representations, followed by bidirectional Mamba for long-range temporal modeling. Experiments on VoiceBank+DEMAND and DNS demonstrate the effectiveness of SFG-Mamba, achieving PESQ scores of 3.57 and 3.52, respectively. Furthermore, additional experiments are designed to investigate the performance drop under increasing temporal heterogeneity. SFG-Mamba achieves the best performance among the compared methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.