BAT: Lightweight Causal Bone-Conduction-Guided Speech Separation with Cross-Modal Conditioning and Adaptive Mask Bounds
Abstract
We study the recovery of wearer and interlocutor speech from a noisy air-conduction (AC) mixture and a synchronized bone-conduction (BC) reference available only from the wearer. BC provides a speaker-specific cue, but its high-frequency attenuation and locally varying reliability complicate its use for separation. We introduce BAT, a compact frame-causal framework that adapts BC conditioning throughout feature fusion and spectral reconstruction. Its Cross-Band Harmonic Reliability Fusion (CHRF) module aggregates cross-band BC context and conditions AC features according to local cross-sensor reliability. The Confidence-Adaptive Mask (CAM) decoder combines separator features with cross-sensor agreement and BC activity to adapt the allowable complex-mask ranges. We also introduce LS-Bone, English AC–BC pairs obtained by controlled re-recording of LibriSpeech, and VibroMix, which combines these pairs with Chinese ABCS speech and environmental noise. On VibroMix-2Mix, BAT achieves 17.3 dB mean SI-SNRi, exceeding the two-output Robust Fusion adaptation by 2.6 dB. Mean SDRi and PESQ improve by 2.5 dB and 0.2, respectively. Under the common causal profiling configuration, BAT uses approximately 46.2% fewer parameters and 34.6% fewer multiply-accumulate operations than the original single-target Robust Fusion model. Code and audio demos are available at https://anonymous.4open.science/w/BAT-Demo-3603/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.