FocusMe: Bone-Conduction-Guided Speech Enhancement in Challenging Acoustic Scenarios
Abstract
Recovering a wearer's speech from an air-conduction (AC) mixture requires both suppressing interference and identifying the target speaker, particularly when the interference includes other voices. We propose FocusMe, a causal speech enhancement system that uses synchronized bone-conduction (BC) speech for two complementary purposes: acoustic reconstruction and enrollment-free speaker conditioning. Cross-Attention Feature Fusion retrieves BC information relevant to the AC mixture and adaptively combines the two signals, while a causal extractor supplies frame-level BC speaker representations to the reconstruction network. The system is trained jointly for speech reconstruction and speaker classification. Across 13 test conditions constructed from ABCS and ESMB, FocusMe achieves the best average perceptual quality, intelligibility, and recognition accuracy among the compared systems, reducing mean character error rate from 22.85% for DCCRN to 17.03%. These results demonstrate the effectiveness of jointly exploiting BC speech for acoustic reconstruction and speaker conditioning under severe noise and competing speech.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.