acceptodds
Under review as a conference paper at ICLR 2027

BiGMNet: Bidirectional Representations with Predictive Supervision for EEG-Guided Target Speaker Extraction

Abstract

Brain-assisted target speaker extraction aims to recover the attended target speech from competing speech, using electroencephalography (EEG) as the sole auxiliary cue. However, architectures relying on local post-fusion mappings and waveform supervision lack explicit coupling between temporal organization and direct supervision of target-speech dynamics, while uniform elementwise updates overlook the distinct roles of internal matrices and other parameters. To address these issues, this paper proposes Bidirectional Griffin Adapter with Multi-Frame Prediction Network (BiGMNet), an EEG-conditioned encoder–mask–decoder framework for target speaker extraction. The method introduces three coordinated components: the Bidirectional Griffin Adapter (BiGriffinAdapter) for bidirectional temporal refinement, direction-aware multi-frame prediction (DA-MFP) for supervising directional intermediate states, and Hybrid Muon–AdamW optimization for parameter-role-aware updates. BiGriffinAdapter combines bidirectional gated recurrence and local attention to refine fused representations for mask estimation, and training-only DA-MFP supervises its pre-concatenation forward and backward states through sequential prediction of future and past continuous clean-speech latents, respectively. Hybrid Muon–AdamW approximately orthogonalizes updates to selected internal matrices and applies elementwise adaptive updates to the remaining parameters according to their computational roles. This paper evaluates subject-specific models on MM-AAD and AVED with ten participants per dataset through EEG-only baseline comparisons and component ablations. BiGMNet achieves 9.76 and 10.90 dB SI-SDR, exceeding IFENet's reported scores by 1.57 and 2.14 dB, respectively, with 12.31 G MACs, 15.5% fewer than IFENet. These results support directly supervising EEG-conditioned temporal states to improve target speaker extraction, with the optimization rule chosen according to parameter roles.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.